Pith. sign in

REVIEW 5 cited by

ReZero is All You Need: Fast Convergence at Large Depth

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.04887 v2 pith:F3ZUAXMO submitted 2020-03-10 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords convergencedeeparchitecturedynamicalfastisometrylayernetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep networks often suffer from vanishing or exploding gradients due to inefficient signal propagation, leading to long training times or convergence difficulties. Various architecture designs, sophisticated residual-style networks, and initialization schemes have been shown to improve deep signal propagation. Recently, Pennington et al. used free probability theory to show that dynamical isometry plays an integral role in efficient deep learning. We show that the simplest architecture change of gating each residual connection using a single zero-initialized parameter satisfies initial dynamical isometry and outperforms more complex approaches. Although much simpler than its predecessors, this gate enables training thousands of fully connected layers with fast convergence and better test performance for ResNets trained on CIFAR-10. We apply this technique to language modeling and find that we can easily train 120-layer Transformers. When applied to 12 layer Transformers, it converges 56% faster on enwiki8.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Controlled Study of Attention-Only Transformers

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Attention-only transformers match standard transformers within 0.006 nats of loss at matched parameter count, with the residual gap localized to low-context parametric recall.

  2. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

    cs.LG 2026-07 unverdicted novelty 6.0 of 10

    SEL weight reparameterization reaches matched OpenWebText validation loss in 1.32–1.49× fewer transformer steps via a sign-aware exponential-linear map and mismatched initialization.

  3. Prototype Transformer: Towards Language Model Architectures Interpretable by Design

    cs.AI 2026-02 conditional novelty 6.0 of 10

    ProtoT is an autoregressive language model whose attention is replaced by learned prototype channels that are claimed to capture nameable concepts and allow targeted edits, at linear sequence cost but with slightly lo...

  4. Residual Matrix Transformers: Scaling the Size of the Residual Stream

    cs.LG 2025-06 conditional novelty 5.0 of 10

    The RMT uses an outer product memory as the residual stream, allowing the residual stream to be scaled up at negligible cost and improving language model efficiency.

  5. NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation

    cs.CL 2025-12 conditional novelty 3.0 of 10

    A gated two-embedding toy model can output near-maximum uncertainty before context and resolve perfectly afterward, but the uncertainty is enforced by a hand-set gate.

Pith tools