REVIEW 5 cited by
ReZero is All You Need: Fast Convergence at Large Depth
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep networks often suffer from vanishing or exploding gradients due to inefficient signal propagation, leading to long training times or convergence difficulties. Various architecture designs, sophisticated residual-style networks, and initialization schemes have been shown to improve deep signal propagation. Recently, Pennington et al. used free probability theory to show that dynamical isometry plays an integral role in efficient deep learning. We show that the simplest architecture change of gating each residual connection using a single zero-initialized parameter satisfies initial dynamical isometry and outperforms more complex approaches. Although much simpler than its predecessors, this gate enables training thousands of fully connected layers with fast convergence and better test performance for ResNets trained on CIFAR-10. We apply this technique to language modeling and find that we can easily train 120-layer Transformers. When applied to 12 layer Transformers, it converges 56% faster on enwiki8.
Forward citations
Cited by 5 Pith papers
-
A Controlled Study of Attention-Only Transformers
Attention-only transformers match standard transformers within 0.006 nats of loss at matched parameter count, with the residual gap localized to low-context parametric recall.
-
Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization
SEL weight reparameterization reaches matched OpenWebText validation loss in 1.32–1.49× fewer transformer steps via a sign-aware exponential-linear map and mismatched initialization.
-
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
ProtoT is an autoregressive language model whose attention is replaced by learned prototype channels that are claimed to capture nameable concepts and allow targeted edits, at linear sequence cost but with slightly lo...
-
Residual Matrix Transformers: Scaling the Size of the Residual Stream
The RMT uses an outer product memory as the residual stream, allowing the residual stream to be scaled up at negligible cost and improving language model efficiency.
-
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
A gated two-embedding toy model can output near-maximum uncertainty before context and resolve perfectly afterward, but the uncertainty is enforced by a hand-set gate.
Discussion (0). Continue with ORCID to comment.