# Claim ledger: hierarchical three-layer Muon/DMFT/BBP paper

This ledger is part of the scientific deliverable.  It prevents exact finite
identities, imported asymptotic theorems, conditional consequences, and
empirical observations from being presented at the same proof level.

## A. Proved in the manuscript

| ID | Statement | Assumptions | Evidence |
|---|---|---|---|
| E1 | The reduced three-layer quadratic network is a polynomial of Gaussian teacher coordinates of total degree at most four. | Fixed finite `k,p,q`; quadratic Hermite activations. | Direct expansion; Proposition `prop:finite-support`. |
| E2 | Its population square loss is exactly the weighted Euclidean distance between multivariate Hermite coefficients. | Gaussian input; target degree at most four. | Hermite orthogonality; Theorem `thm:exact-closure`. |
| E3 | Population gradient flow closes exactly through the coefficient Jacobian, and the Hessian splits into Gauss--Newton and residual curvature. | Twice differentiable parameterization. | Chain rule; Theorem `thm:exact-closure`. |
| E4 | The degree-2/3/4 Schur matrices are residualized coefficient-Jacobian Gram matrices and are positive semidefinite. | Moore--Penrose inverses; fixed state. | Orthogonal projection identity; Proposition `prop:schur-hierarchy`. |
| E5 | The projected total Hessian is bounded below by the active Schur curvature minus the projected residual-curvature norm. | Fixed state. | Weyl/Rayleigh bound; Proposition `prop:active-curvature`. |
| E6 | Every nonzero Muon-\(a\) matrix update used in the experiments is a strict first-order descent direction for \(a\geq0\); a standard smoothness bound gives a sufficient step size. | Nonzero gradient; locally Lipschitz gradient. | Singular-value calculation; Proposition `prop:muon-descent`. |
| E7 | The analytic preactivation gradient and Hessian of the link have the block formulas stated in the paper. | Quadratic Hermite architecture. | Direct differentiation; Proposition `prop:link-hessian` and automatic-differentiation test. |
| E8 | At a reduced state embedded in ambient dimension and evaluated on a fresh sample, the orthogonal direct-weight Hessian is exactly matrix-weighted block Wishart conditional on the teacher coordinates. | Fresh Gaussian sample; fixed `k,p,q`; weights embedded in teacher span. | Independence and Kronecker calculus; Theorem `thm:fresh-block-wishart`. |
| E9 | The remaining teacher-coordinate and finite-parameter sector obeys an exact finite-dimensional Schur determinant outside the orthogonal spectrum. | Spectral parameter outside bulk spectrum. | Block determinant identity; Proposition `prop:finite-schur`. |
| E10 | Spectral clipping obeys the deterministic operator-norm perturbation bound stated in the paper. | Self-adjoint blocks. | Triangle property of the Kronecker norm; Lemma `lem:clipping`. |
| E11 | At fixed row dimension, a singular-value power update is an exact finite functional of the row gradient Gram matrix. | Full row rank, or Moore--Penrose calculus on the positive singular subspace. | SVD functional calculus; Proposition `prop:finite-muon-action`. |
| E12 | Hermite-tail norms control truncation errors in population risk, gradient, and Hessian by the explicit bounds in the paper. | Finite parameter dimension; the coefficient map into Gaussian \(L^2\) is twice Fréchet differentiable. | Orthogonal projection and Cauchy--Schwarz; Theorem `thm:hermite-tail-transfer`. |
| E13 | An isolated finite population-Hessian cluster is stable when the Hermite-tail Hessian perturbation is smaller than half its gap. | E12 and the stated spectral gap. | Weyl perturbation inequality; Corollary `cor:spectral-tail-stability`. |
| E14 | The full-rank Muon-\(a\) spectral channel has the stated exact Fréchet derivative, including singular-vector rotation and unit-RMS normalization. | Full row rank; numerical floor inactive; spectral matrix functional calculus. | Daleckii--Krein formula and chain rule; Proposition `prop:muon-frechet`. |

## B. Imported theorems, applied only under their hypotheses

| ID | Imported result | Use in this manuscript | Boundary |
|---|---|---|---|
| I1 | Montanari--Saeed block-Wishart global law and variational support-edge formula. | Conditional limiting law for clipped fresh orthogonal Hessian blocks. | Does not supply the anisotropic local law or finite-rank outlier theorem. |
| I2 | Dandi--Pesce--Zdeborová--Krzakala controlled hierarchical feature recovery. | Motivation for effective-dimension reduction and staged controls. | Does not prove joint Muon training. |
| I3 | Nichani--Damian--Lee and Wang--Nichani--Lee three-layer feature-learning guarantees. | Comparison to layerwise/frozen schedules and \(d^k\)-type feature recovery. | Their algorithms and architecture assumptions differ from the joint quadratic residual model. |
| I4 | Wortsman-Zurich--Tabanelli--Dandi--Krzakala--Loureiro sharp sequential spectral thresholds. | Prediction that a cascade of feature-wise BBP transitions aggregates into a scaling law. | The present fixed-rank pilot cannot establish the growing-rank exponent. |
| I5 | Braun--Loureiro--Quang Minh--Imaizumi deterministic Volterra/Stieltjes phases. | Population anisotropy benchmark and three-phase interpretation. | Phase retrieval is not the present three-layer network. |
| I6 | Kramp--Lindner--Helias random-feature DMFT. | Frozen-feature control and notation for response/noise kernels. | The present features move, so extra coupled kernels are required. |
| I7 | Paquette et al. spectral-optimizer dynamics. | Motivation for SignSVD/Muon and power filters. | No universal exponent is imported into the three-layer model. |
| I8 | Berthier--Montanari--Nguyen state evolution for nonseparable AMP denoisers. | Evidence that a nonseparable spectral channel can admit state evolution under Gaussian-design hypotheses. | Does not cover the hierarchical reused-data factor graph or Muon without Onsager correction. |
| I9 | Rangan--Schniter--Fletcher VAMP. | Orthogonally invariant linear-operator benchmark for spectral message passing. | The present operators and features evolve and are nonlinear. |
| I10 | Lesieur--Krzakala--Zdeborová Low-RAMP/TAP/replica correspondence. | Template for relating AMP fixed points, Onsager reactions, and replica-symmetric potentials. | Their low-rank observation models differ from joint three-layer training. |
| I11 | Coeurdoux--Ferré--Bouchaud transient-BBP gradient-flow solution. | Minimal exactly soluble control for absent, persistent, and transient spectral branches. | Their observable is a symmetrized linear-regression weight matrix with a two-block Wigner-type bath, not the three-layer Hessian or a block-Wishart ensemble. |
| I12 | Maillard--Ben Arous--Biroli, Asgari--Montanari--Saeed, Montanari--Saeed, and Maillard--Bonnaire--Biroli Kac--Rice formulas. | Landscape-complexity and conditioned-Hessian interface; motivation for a BBP--Kac--Rice calculation. | They do not give the critical-point complexity of the present parameterization, and stationary conditioning does not determine transient Muon dynamics. |
| I13 | Classical and discontinuous BBP theory. | Criteria to distinguish regular continuous detachment from a jump in eigenvector overlap at a nonstandard edge. | The edge exponent and finite-Schur contact order have not been derived for the three-layer Hessian. |
| I14 | Giammanco--Valigi--Cammarota sparse non-Hermitian cavity and population dynamics. | Computational template for distribution-valued resolvent messages and a warning about strong-disorder support estimation. | Their locally tree-like non-Hermitian ensemble is not the dense Hermitian fresh Hessian. |
| I15 | Montanari--Saeed universality of empirical-risk minima. | Gaussian-replacement benchmark for fixed-index optimization under delocalization. | Does not establish universality of trajectories, Hessian edges, BBP residues, or optimizer response kernels. |

## C. Conditional or conjectural statements

| ID | Statement | Missing proof |
|---|---|---|
| C1 | Joint reused-sample training admits the stated three-layer DMFT with two-time covariance and response kernels for all trainable blocks. | Dynamic cavity/leave-one-out derivation and well-posedness. |
| C2 | The exact finite Schur determinant converges to a deterministic BBP equation along the training path. | Anisotropic local law uniform in time and control of finite-signal resolvent entries. |
| C3 | Raw, unclipped polynomial Hessian extremes obey the same edge/outlier description after an explicit tail renormalization. | Heavy-tail truncation removal and extreme-block analysis. |
| C4 | A block-specific Muon policy \(a_V(t),a_U(t),a_A(t)\) is asymptotically optimal. | Controlled DMFT/adjoint analysis for the three coupled matrices. |
| C5 | Growing latent rank with power-law feature strengths produces the feature-wise thresholds and aggregate risk exponent proposed in the paper. | Joint \(d,n,k\) limit for the hierarchical network. |
| C6 | The DMFT state determines moving RMT edges and simple BBP branches under the four hypotheses of Proposition `prop:conditional-dynamic-closure`. | The implication is proved, but its DMFT, time-uniform MDE/local-law, finite-Schur, and tail/data-reuse hypotheses are not established here. |
| C7 | A blockwise long-memory spectral-AMP algorithm with Muon-\(a\) denoisers closes on the proposed state and matches stable replica-symmetric fixed points. | Factor-graph derivation, matrix Onsager kernels, nonseparable state evolution, and replica stability. |
| C8 | The present three-layer empirical loss admits a finite-dimensional annealed and quenched Kac--Rice complexity indexed by risk, teacher overlaps, and representation invariants. | Gauge fixing, nondegenerate gradient conditioning, conditioned block-Hessian law, exponential determinant asymptotics, and replica/second-moment control. |
| C9 | A moving three-layer branch can be classified as continuous, discontinuous, or transient from its limiting contact. | Regular edge expansion, contact order of the Schur determinant, residue asymptotics, and finite-size scaling. |

## D. Empirical statements

The checked-in experiments support only the following empirical claims:

1. In the pre-specified twenty-pair confirmation, polar Muon, \(a=1/7\), and
   \(a=1/3\) beat gradient descent on `9/20`, `8/20`, and `10/20` pairs.
   Every paired bootstrap interval for the median log-risk difference contains
   zero; there is no demonstrated endpoint dominance.
2. The confirmatory Kaplan--Meier degree-2 half-error median is `160` steps for
   GD and `80` for all three spectral rules.  Degree-4 medians are `480` (GD),
   `240` (polar), `240` (\(a=1/7\)), and `320` (\(a=1/3\)), but terminal event
   coverage is not uniformly improved.
3. On the exploratory three-pair pilot, polar Muon and both transferred powers beat gradient
   descent on two runs, but none dominates on every pair.
4. Median endpoint exact population risk is `0.08712` for polar Muon and
   `0.7569` for gradient descent in the main pilot.
5. Degree-4 error is the late bottleneck in the displayed pilot repeat; polar Muon
   reaches its half-error threshold at step 680 while gradient descent does not
   reach it by step 1200.
6. At \(n=2048\) in the five-repeat nested-sample panel, the median risks are
   `0.2592` (GD), `0.1038` (polar), `0.06948` (`a=1/7`), and `0.08954`
   (`a=1/3`).
7. Raw informative eigenmodes move relative to the fresh block-Wishart bulk,
   but the current data do not establish a monotone or asymptotically sharp BBP
   crossing time.
8. Clipping a small fraction of polynomial blocks can have a large
   operator-norm effect; raw and bounded-theorem spectra must remain separate.
9. The dense coupled panel places exact sector errors, exact finite Schur
   margins, and fresh finite-size spectra on one training clock.  Temporal
   proximity is descriptive and is not evidence of a causal DMFT-to-BBP
   mechanism.

The journal manuscript must not convert any item in B, C, or D into an
unqualified theorem.
