# Literature map and proof interfaces

This file records why each research lineage appears in the manuscript.  It is
not a citation dump: a paper enters the main text only when it supplies a
theorem, a solvable control, or a clearly delimited open interface.

The most important distinction is between four laws that are often conflated:

| object | conditioning | method | status here |
|---|---|---|---|
| population Hessian | prescribed finite parameter | exact Hermite algebra | proved |
| fresh empirical Hessian | trained state fixed, new Gaussian observations | conditional block-Wishart RMT | exact identity; bounded global law imported |
| same-data trajectory Hessian | observations reused to produce the state and Hessian | DMFT plus dynamic leave-one-out/local law | open |
| critical-point Hessian | gradient zero, risk/overlaps/invariants fixed | Kac--Rice plus conditioned RMT | open for this network |

## 1. Depth and sequential feature recovery

- [Dandi, Pesce, Zdeborová, and Krzakala](https://arxiv.org/abs/2502.13961)
  establish a controlled depth advantage through staged effective-dimension
  reduction.  Their schedule is a control, not a proof for joint Muon.
- [Nichani, Damian, and Lee](https://arxiv.org/abs/2305.06986) and
  [Wang, Nichani, and Lee](https://arxiv.org/abs/2311.13774) provide
  three-layer hierarchical feature-learning guarantees under different
  architectures and training schedules.
- [Wortsman-Zurich et al.](https://arxiv.org/abs/2605.14567) turn sequential
  spectral thresholds into a solvable scaling law.  The fixed-rank experiment
  in this repository cannot identify their growing-rank exponent.
- [Pillaud-Vivien and Schertzer](https://arxiv.org/abs/2505.21336) and
  [Braun et al.](https://arxiv.org/abs/2511.18661) supply single-index
  population-dynamics controls.  The latter separates fast escape,
  macroscopic convergence, and spectral-tail learning under power-law
  covariance.

## 2. Field theory, DMFT, and rich versus lazy learning

- [Helias and Dahmen](https://arxiv.org/abs/1901.10416) give the
  MSRJD/effective-action framework used to state the dynamic target.
- [Segadlo et al.](https://arxiv.org/abs/2112.05589) derive Gaussian-process
  saddles and finite-width corrections for deep and recurrent networks.
- [Kramp, Lindner, and Helias](https://arxiv.org/abs/2602.23039) solve a
  fixed-feature/random-feature regression DMFT.  It is the lazy
  NNGP/NTK/random-feature control: freezing the representation must recover
  its response and colored-noise structure.
- [van Meegen and Sompolinsky](https://arxiv.org/abs/2406.16689) provide a
  static non-lazy representation-learning benchmark; activation-dependent
  coding cannot be replaced by a kernel-only closure.
- [Dick, van Meegen, and Helias](https://arxiv.org/abs/2309.14973) motivate
  the effective-action/renormalized-field-theory language.  The manuscript
  does not claim an RG fixed point or universal critical exponent.

## 3. Block-Wishart bulk and local laws

- [Montanari and Saeed](https://arxiv.org/abs/2606.27774) prove the global
  law, matrix Stieltjes equation, variational support edges, and logarithmic
  potential for bounded fixed-block Wishart matrices.  This theorem is
  imported only after spectral clipping of the fresh conditional blocks.
- [Knowles and Yin](https://arxiv.org/abs/1410.3516) represent the
  anisotropic-local-law level of control required by finite teacher-sector
  resolvent quadratic forms.  Their theorem is not directly applied to the
  moving block ensemble.
- [Montanari and Saeed](https://arxiv.org/abs/2202.08832) prove universality
  of fixed-index empirical-risk minima under a delocalization condition.
  This does not imply trajectory, Hessian-edge, residue, or Muon-response
  universality.

## 4. BBP and time-dependent spectral branches

- [Baik, Ben Arous, and Péché](https://doi.org/10.1214/009117905000000233)
  is the classical sample-covariance transition.
- [Benaych-Georges and Nadakuditi](https://arxiv.org/abs/0910.2120) give a
  general finite-rank deformation framework for outlier eigenvalues and
  eigenvectors.
- [Bonnaire, Biroli, and Cammarota](https://arxiv.org/abs/2403.02418) study
  the time-dependent Hessian in phase retrieval and a finite-time loss of an
  informative negative direction.
- [Annesi, Bocchi, and Cammarota](https://arxiv.org/abs/2510.18435) analyze
  continuous and discontinuous BBP behavior in a simple overparameterized
  neural model at initialization.
- [Bocchi et al.](https://arxiv.org/abs/2604.27992) show that sufficiently
  nonstandard edge vanishing can produce a discontinuous overlap jump and a
  large finite-size precursor regime.
- [Coeurdoux, Ferré, and Bouchaud](https://arxiv.org/abs/2604.18450) provide
  the minimal exactly soluble dynamic control: explicit linear gradient flow,
  a two-block Wigner-type bath, a \(2\times2\) Dyson equation, and a rank-two
  determinant yielding absent, persistent, or transient branches.

The last paper is a methodological control, not a special case of the present
Hessian: its spectral observable is a symmetrized trained weight matrix,
whereas the fresh three-layer Hessian has a block-Wishart transverse bath.

## 5. Kac--Rice and landscape topology

- [Kac](https://doi.org/10.1090/S0002-9904-1943-08069-X),
  [Rice](https://doi.org/10.1002/j.1538-7305.1945.tb00453.x), and
  [Adler and Taylor](https://doi.org/10.1007/978-0-387-48116-6) provide the
  exact counting identity and geometric regularity framework.
- [Auffinger, Ben Arous, and Černý](https://arxiv.org/abs/1003.1129) connect
  spin-glass complexity to conditioned GOE spectra.
- [Maillard, Ben Arous, and Biroli](https://arxiv.org/abs/1912.02143) extend
  Kac--Rice to non-Gaussian empirical-risk landscapes and distinguish
  annealed from replicated quenched complexity.
- [Asgari, Montanari, and Saeed](https://arxiv.org/abs/2502.01953) give
  general high-dimensional local-minimum formulas under explicit hypotheses.
- [Montanari and Saeed](https://arxiv.org/abs/2602.14969) prove an annealed
  rate formula and sufficient conditions for topological rate
  trivialization.
- [Maillard, Bonnaire, and Biroli](https://arxiv.org/abs/2602.17779) combine
  energy/overlap-conditioned Kac--Rice with a BBP analysis of the Hessian in
  phase retrieval.

A three-layer Kac--Rice theorem still needs gauge fixing, a nondegenerate
conditioned gradient law, the Hessian distribution conditional on all
layerwise overlaps and representation invariants, exponential determinant
asymptotics, and annealed/quenched control.  Even then it would describe
stationary topology, not the causal transient law of deterministic Muon.

## 6. Cavity, replica, AMP, and Muon

- [Bayati and Montanari](https://arxiv.org/abs/1001.3448) establish scalar
  state evolution for dense Gaussian AMP.
- [Berthier, Montanari, and Nguyen](https://arxiv.org/abs/1708.03950) extend
  state evolution to nonseparable Lipschitz denoisers.
- [Rangan, Schniter, and Fletcher](https://arxiv.org/abs/1610.03082) provide
  the VAMP benchmark for right-orthogonally invariant linear operators.
- [Lesieur, Krzakala, and Zdeborová](https://arxiv.org/abs/1701.00858)
  connect low-rank AMP, TAP equations, and replica-symmetric potentials.
- [Bolthausen](https://arxiv.org/abs/1201.2891) gives an iterative
  construction of TAP solutions in the SK model.
- [Rogers et al.](https://arxiv.org/abs/0803.1553) and
  [Neri and Metz](https://arxiv.org/abs/1205.0702) develop cavity spectral
  recursions for sparse matrices; their
  [outlier analysis](https://doi.org/10.1103/PhysRevLett.117.224101)
  treats locally tree-like non-Hermitian ensembles.
- [Giammanco, Valigi, and Cammarota](https://arxiv.org/abs/2606.25925) use
  population dynamics for a sparse non-Hermitian Jacobian-like ensemble and
  explicitly expose a strong-disorder support-estimation failure.
- [Paquette et al.](https://arxiv.org/abs/2605.09552) supply the controlled
  spectral-optimizer benchmark for Muon/SignSVD.

Muon-\(a\) is called **AMP-like only as a nonseparable spectral channel**.
The tested optimizer is not AMP: it has no Onsager reaction.  The manuscript
derives the exact full-rank Fréchet derivative of the normalized spectral
map because that derivative is the input to a future matrix-valued,
long-memory Onsager kernel.

## Remaining intersections

The unresolved research program is the intersection, not the union, of these
literatures:

1. derive a reused-sample MSRJD or dynamic-cavity closure for all trainable
   blocks;
2. derive the matrix Onsager reaction for the Muon spectral channel;
3. prove time-uniform block-Wishart global and anisotropic local laws;
4. remove polynomial clipping or solve the extreme-observation process;
5. prove convergence and contact order of the finite Schur determinant;
6. distinguish continuous, discontinuous, and transient contacts;
7. derive the energy/overlap-conditioned three-layer Kac--Rice law;
8. take the growing-rank limit before asserting a universal scaling exponent.
