Pith. sign in

REVIEW 4 cited by

Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01814 v2 pith:PMQNFFQI submitted 2023-08-03 cs.LG cond-mat.dis-nncs.NEmath.PR

classification cs.LGcond-mat.dis-nncs.NEmath.PR
keywords tensoradaptiveoptimizersprogramsadamkernelneuralresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Going beyond stochastic gradient descent (SGD), what new phenomena emerge in wide neural networks trained by adaptive optimizers like Adam? Here we show: The same dichotomy between feature learning and kernel behaviors (as in SGD) holds for general optimizers as well, including Adam -- albeit with a nonlinear notion of "kernel." We derive the corresponding "neural tangent" and "maximal update" limits for any architecture. Two foundational advances underlie the above results: 1) A new Tensor Program language, NEXORT, that can express how adaptive optimizers process gradients into updates. 2) The introduction of bra-ket notation to drastically simplify expressions and calculations in Tensor Programs. This work summarizes and generalizes all previous results in the Tensor Programs series of papers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Rank survival in Transformer blocks is governed by a branch-to-skip ratio law (βα^M√L), a mean-spike coherence c_ℓ=E[σ]²/E[σ²], and a Marchenko–Pastur width threshold m/d=1/p(σ).

  2. Dimension-adapted Momentum Outscales SGD

    stat.ML 2025-05 conditional novelty 7.0 of 10

    DANA, with dimension- and time-dependent momentum, provably outscales SGD on power-law random features when 2α>1, improving loss exponents and compute-optimal curves.

  3. Distillation Scaling Laws

    cs.LG 2025-02 conditional novelty 7.0 of 10

    A distillation scaling law predicts student cross-entropy from teacher loss, student size, and data, and gives compute-optimal teacher-student allocations.

  4. SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SOAP and Muon, stabilized by per-step QR eigenbasis updates and KL-Shampoo covariance accumulation, beat AdamW on large-batch LLM pretraining up to 100M-token batches.

Pith tools