Pith. sign in

REVIEW 8 cited by

A Spectral Condition for Feature Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17813 v2 pith:KP2GUVDW submitted 2023-10-26 cs.LG

classification cs.LG
keywords featurelearningspectralnetworknetworksneuralnormscaling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network's internal representations evolve nontrivially at all widths, a process known as feature learning. Here, we show that feature learning is achieved by scaling the spectral norm of weight matrices and their updates like $\sqrt{\texttt{fan-out}/\texttt{fan-in}}$, in contrast to widely used but heuristic scalings based on Frobenius norm and entry size. Our spectral scaling analysis also leads to an elementary derivation of \emph{maximal update parametrization}. All in all, we aim to provide the reader with a solid conceptual understanding of feature learning in neural networks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Algorithmic Separation between Constant-Depth and Logarithmic-Depth Neural Networks

    cs.LG 2026-07 accept novelty 8.0 of 10

    Logarithmic-depth networks trained by layerwise coordinate descent can learn hierarchical staircase Boolean functions that constant-depth networks with bounded spectral norms cannot approximate.

  2. Conditional Optimal Bridge for Riemannian Activation Steering

    cs.LG 2026-07 accept novelty 7.0 of 10

    Casting activation steering as a Schrödinger Bridge on the residual hypersphere derives the log-density-ratio objective and yields query-adaptive directions that beat fixed baselines without OOD collapse.

  3. Training Transformers with Enforced Lipschitz Constants

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Transformers can be trained with enforced spectral-norm constraints throughout training, but competitive accuracy requires an astronomical Lipschitz upper bound.

  4. LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Lipschitz-constrained SSD variants improve white-box adversarial robustness in an attack-agnostic way and remain complementary to adversarial training on VOC, KITTI, and LARD.

  5. Customizing the Inductive Biases of Softmax Attention using Structured Matrices

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Structured-matrix scoring functions, BTT and MLR, let attention escape the low-rank bottleneck and add a distance-dependent compute bias, improving accuracy for fixed compute on regression, language modeling, and forecasting.

  6. Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.

  7. Low-rank Momentum Factorization for Memory Efficient Training

    cs.LG 2025-07 reject novelty 6.0 of 10

    MoFaSGD keeps a low-rank factored momentum and uses its singular vectors as the update direction, achieving LoRA-level memory with competitive fine-tuning performance, but its convergence proof is flawed.

  8. Scale Weight Decay and Train Better

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Muon with weight decay scaled by η/η_max reaches the same MoE validation loss ~30% faster than constant-decay Muon while preserving asymptotic stationarity of the unregularized objective.

Pith tools