Hierarchical three-layer learning

Three Layers, One Hierarchy: From Exact Hermite Dynamics to DMFT and BBP

Theory first, experiment second. This article builds the deterministic and random-matrix bridge. The companion article, Does Muon Help a Hierarchical Three-Layer Model?, gives the finite-sample evidence and the complete numerical caveats.

The question

Depth is useful when one layer discovers a representation that makes the next learning problem simpler. That slogan is compelling, but it hides three different mathematical questions:

  1. Can a finite model exhibit a genuine hierarchy of features rather than merely contain three parameter matrices?
  2. Does a spectral optimizer such as Muon change the time at which successive feature sectors become accessible?
  3. When finite data add a random Hessian bulk, when does a learned direction become visible as an isolated eigenvector?

The model studied here is deliberately small enough that its population loss, gradient, and Hessian can be calculated exactly. This gives a deterministic baseline before introducing finite-sample training, dynamical mean-field theory (DMFT), or a Baik–Ben Arous–Péché (BBP) transition.

An exact three-layer hierarchy

Let $z\in\mathbb R^k$ denote the teacher coordinates of a Gaussian input and let

\[\begin{aligned} h_1 &= Vz+b, & q &= \operatorname{He}_2(h_1)=h_1^{\odot 2}-1,\\ h_2 &= Uz+Aq+c, & s &= \operatorname{He}_2(h_2)=h_2^{\odot 2}-1,\\ f_\theta(z)&=a_0+\alpha^\top q+\beta^\top s. \end{aligned}\]

The first block creates quadratic features. The second mixes the original coordinates with those features and applies another quadratic transformation. The readout combines both levels. Consequently, the prediction contains Hermite sectors through order four.

This is a three-trainable-layer model in the convention of the hierarchical-feature literature. The experiments use the untied version $U\neq V$, because tying the two matrices makes the reduced dynamics substantially stiffer.

Why the population calculation is deterministic

Expand the prediction and teacher in the multivariate Hermite basis. If $c(\theta)$ is the finite vector of coefficients through order four, $c_\star$ the teacher vector, and $D_{\gamma\gamma}=\gamma!$, then the population risk is

\[\mathcal R(\theta)=\frac12\big(c(\theta)-c_\star\big)^\top D\big(c(\theta)-c_\star\big).\]

Writing $J_c$ for the coefficient Jacobian gives

\[\nabla\mathcal R=J_c^\top D(c-c_\star),\]

and

\[\nabla^2\mathcal R =J_c^\top D J_c +\sum_\gamma D_{\gamma\gamma}(c_\gamma-c_{\star,\gamma}) \nabla^2 c_\gamma.\]

The first term is Gauss–Newton curvature. The second is residual curvature and can be indefinite away from interpolation. Gauss–Hermite quadrature evaluates all polynomial expectations exactly here, so this is not a Monte Carlo approximation and not yet a DMFT.

Splitting $c-c_\star$ by Hermite order gives errors $E_2,E_3,E_4$. Their weighted Gram matrices and Schur complements $S_2,S_3,S_4$ quantify which new directions remain after lower-order sectors have been accounted for. They are finite-model observables of hierarchical coarse-graining.

How this connects to the depth literature

The physical picture is a sequence of effective dimension reductions:

ambient Gaussian input
        ↓
first recovered teacher subspace
        ↓
nonlinear lifted features
        ↓
smaller effective problem for the next layer
        ↓
final readout

Dandi, Pesce, Zdeborová, and Krzakala formalize this mechanism for hierarchical targets with decreasing effective dimensions. Nichani, Damian, and Lee and Wang, Nichani, and Lee prove advantages for staged three-layer procedures that first recover a lower-degree feature and then learn an outer function. Their assumptions use layerwise schedules, freezing, or sample splitting; they do not directly prove the joint backpropagation dynamics considered here.

Wortsman-Zurich, Tabanelli, Dandi, Krzakala, and Loureiro derive scaling laws from sequential spectral feature recovery. In their growing hierarchy, feature $i$ becomes visible at a sharp sample scale and the aggregate of many transitions becomes a smooth power law. Our fixed $k=3$ system tests the ordering mechanism, not their growing-rank exponent.

Pillaud-Vivien and Schertzer show that a direction and its nonlinear link can be learned jointly in a Gaussian single-index model. This supports allowing the readout to adapt, but their alternating single-index theorem is not a theorem for the coupled matrices $V,U,A$.

What Muon changes

For a matrix gradient $G=P\operatorname{diag}(\sigma_i)Q^\top$, the tested spectral family is

\[\Phi_a(G)=P\operatorname{diag}\!\left[ \left(\frac{\sigma_i}{\overline\sigma}\right)^a \right]Q^\top,\]

followed by unit-RMS normalization. The case $a=0$ is exact polar, or SignSVD, Muon: all nonzero singular directions receive equal spectral magnitude. Only the matrix blocks $V,U,A$ receive this transformation; vectors and biases use ordinary gradient descent.

The empirical panel also includes $a=1/7$ and $a=1/3$. These are transfer controls from the separate phase-retrieval/Volterra program. They are not derived optima for this hierarchical model. A genuine theory should allow different, state-dependent powers and clocks for $V$, $U$, and $A$.

The DMFT target, and why it is still open

The deterministic Hermite system contains no disorder-averaged training noise. A full reused-data DMFT would need, for every matrix layer ℓ, two-time correlations and responses such as

\[C_\ell(t,s)=\frac{1}{d_\ell}\langle W_\ell(t),W_\ell(s)\rangle, \qquad R_\ell(t,s)=\frac{\delta W_\ell(t)}{\delta \eta_\ell(s)},\]

as well as cross-layer kernels because $A$ transports first-layer features into the second layer. The effective process should contain deterministic drift, a memory integral, and self-consistent noise. Muon adds the Fréchet derivative of the polar or matrix-power map to the response.

The frozen-random-feature DMFT of Kramp, Lindner, and Helias is the correct benchmark for mode-wise scaling and renormalized response. It cannot simply be copied: in the present model the features and their spectrum move with $V,U,A$.

A controlled route would start with fresh batches, then prove a finite-time cavity or leave-one-out transfer to reused samples, and only then derive the corresponding Onsager response for Muon.

From a finite Hessian to a block-Wishart bulk

At a fixed reduced state, embed $z$ into an ambient Gaussian input $x=(z,\xi)\in\mathbb R^d$. Let $h=(h_1,h_U)$ collect the direct preactivations and let Δ be the sample residual. The per-sample preactivation Hessian is

\[T(z)=\nabla_h f\,\nabla_h f^\top+\Delta\,\nabla_h^2 f.\]

The extensive Hessian with respect to the direct weight blocks is exactly

\[H_W=\frac1n\sum_{\mu=1}^n T(z_\mu)\otimes x_\mu x_\mu^\top.\]

On the orthogonal coordinates,

\[H_\perp=\frac1n\sum_{\mu=1}^n T(z_\mu)\otimes \xi_\mu\xi_\mu^\top.\]

For a fresh sample, $T(z_\mu)$ and $\xi_\mu$ are independent. The orthogonal Hessian is therefore exactly a matrix-weighted block-Wishart matrix. Montanari and Saeed characterize its limiting bulk through the matrix transform

\[K(S)=\mathbb E\!\left[T(I+ST)^{-1}\right]-\alpha^{-1}S^{-1},\]

with the matrix Stieltjes transform obtained from $K(S)=zI$. The remaining teacher coordinates and finite parameters form a finite signal sector. In principle, a Schur determinant then locates detached eigenvalues and its residue identifies whether their eigenvectors align with the teacher.

There is one important complication: polynomial links make $T(z)$ unbounded. The theorem-compatible calculation therefore clips the block spectrum and reports it separately from the raw polynomial Hessian. A small clipping fraction can still have a large operator-norm effect near the extreme eigenvalues.

What is exact and what is not

Statement Status
Hermite coefficient map, population loss, gradient, and Hessian exact finite algebra
orderwise errors and Schur complements exact finite algebra
fresh orthogonal Hessian is block-Wishart exact conditional identity
variational bulk edges for clipped fresh blocks imported theorem under bounded-block assumptions
joint reused-data DMFT for $V,U,A$ open
universal Muon exponent for all three layers not established
raw polynomial BBP outlier theorem open tail/local-law problem
same-data Hessian along the training trajectory open leave-one-out problem

Takeaway

The deterministic model gives a clean hierarchy: lower Hermite sectors create the representation on which higher sectors depend. Muon changes the geometry of the matrix update, while finite samples add a block-Wishart spectral background. These pieces fit into one physical story, but they live at different proof levels. The exact Hermite reduction is the foundation; DMFT and BBP are the next mathematical bridges, not labels to attach prematurely.

Code, report, and companion article