Pith. sign in

REVIEW 5 cited by

Models of Heavy-Tailed Mechanistic Universality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.03470 v1 pith:JJTTVEM7 submitted 2025-06-04 stat.ML cs.LG

Models of Heavy-Tailed Mechanistic Universality

classification stat.ML cs.LG
keywords heavy-tailedmodelneurallawsbehaviorcorrelationdeepdensities
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent theoretical and empirical successes in deep learning, including the celebrated neural scaling laws, are punctuated by the observation that many objects of interest tend to exhibit some form of heavy-tailed or power law behavior. In particular, the prevalence of heavy-tailed spectral densities in Jacobians, Hessians, and weight matrices has led to the introduction of the concept of heavy-tailed mechanistic universality (HT-MU). Multiple lines of empirical evidence suggest a robust correlation between heavy-tailed metrics and model performance, indicating that HT-MU may be a fundamental aspect of deep learning efficacy. Here, we propose a general family of random matrix models -- the high-temperature Marchenko-Pastur (HTMP) ensemble -- to explore attributes that give rise to heavy-tailed behavior in trained neural networks. Under this model, spectral densities with power laws on (upper and lower) tails arise through a combination of three independent factors (complex correlation structures in the data; reduced temperatures during training; and reduced eigenvector entropy), appearing as an implicit bias in the model structure, and they can be controlled with an "eigenvalue repulsion" parameter. Implications of our model on other appearances of heavy tails, including neural scaling laws, optimizer trajectories, and the five-plus-one phases of neural network training, are discussed.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Depth, Not Data: An Analysis of Hessian Spectral Bifurcation

    cs.LG 2026-01 unverdicted novelty 7.0

    Deep linear networks with balanced data covariance exhibit Hessian spectral bifurcation whose dominant-to-bulk eigenvalue ratio scales linearly with depth.

  2. AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization

    cs.LG 2026-06 unverdicted novelty 6.0

    AlphaQ performs calibration-free mixed-precision quantization of MoE models by allocating higher bits to experts whose weight spectra exhibit stronger heavy-tailed structure according to HT-SR theory, outperforming ca...

  3. Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

    cs.LG 2025-09 conditional novelty 6.0

    For diagonal and quadratic two-layer networks, training maps to LASSO and matrix compressed sensing, yielding a full phase diagram of excess-risk scaling exponents and a spectral characterization of the trained weights.

  4. HTMuon: Improving Muon via Heavy-Tailed Spectral Correction

    cs.LG 2026-03 unverdicted novelty 5.0

    HTMuon modifies Muon to produce heavier-tailed updates and weight spectra via HT-SR theory, yielding up to 0.98 lower perplexity on LLaMA pretraining and serving as a plug-in for other Muon variants.

  5. Depth, Not Data: An Analysis of Hessian Spectral Bifurcation

    cs.LG 2026-01 conditional novelty 5.0

    Even with perfectly balanced and whitened data, a deep linear network's Hessian exhibits a two-cluster spectrum whose dominant-to-bulk eigenvalue ratio grows linearly with depth.