Geometric singular learning
Dead Directions: Geometric Singular Learning · arXiv:2606.059571
This is the hub every other paper hangs off. It starts from the two traditions that circled the same degeneracy for decades without meeting, Watanabe's singular learning theory, built for the singular regime, and Amari's information geometry, sharp everywhere else, and identifies the object that lives on the edge between them: the dead direction, a direction in weight space whose Fisher information collapses at a measurable rate. The central theorem is the trajectory-rate translation: the KL order $k$ of a singularity, which classically takes an algebraic-geometry resolution to extract, is the slope $2(k-1)$ of a log–log Fisher decay, something an ordinary training run can expose. From that primitive the paper builds outward: the per-layer K-FAC bridge with its forward–backward duality, the composition rules that lift one layer's reading to a deep network, and the quotient theorem that says exactly which optimizers leave the rate legible.
It is a theory paper in the strict sense: definitions, theorems, proofs, and a closing register of what stays open, with the empirical work delegated to the spokes.
Where it is explained here
- The object and its order: what is a dead direction#3
- The central theorem, slope $2(k-1)$: reading the order of a dead direction#4
- Why the order matters, $\lambda = 1/(2k)$: why singular models generalize#5
- The per-layer bridge and A–G duality: reading the order one layer at a time#6
- The quotient theorem's setting: why the dead direction depends on the optimizer#12 and the arc that follows it
The signature demo
The bridge itself, one integer moving two exponents:
This is the hub the other four hang off: each of them takes one of its claims out to real networks.
- Tejas Pradeep Shirodkar, Dead Directions: Geometric Singular Learning, arXiv:2606.05957 (2026). ↩︎