Pith. sign in

REVIEW 1 cited by

B2T Connection: Serving Stability and Performance in Deep Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.00330 v2 pith:HPRJRRSW submitted 2022-06-01 cs.LG cs.CL

classification cs.LGcs.CL
keywords post-lnpre-lntrainingtransformersdeeplayersconnectioneffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

From the perspective of the layer normalization (LN) positions, the architectures of Transformers can be categorized into two types: Post-LN and Pre-LN. Recent Transformers tend to be Pre-LN because, in Post-LN with deep Transformers (e.g., those with ten or more layers), the training is often unstable, resulting in useless models. However, Post-LN has consistently achieved better performance than Pre-LN in relatively shallow Transformers (e.g., those with six or fewer layers). This study first investigates the reason for these discrepant observations empirically and theoretically and made the following discoveries: 1, the LN in Post-LN is the main source of the vanishing gradient problem that leads to unstable training, whereas Pre-LN prevents it, and 2, Post-LN tends to preserve larger gradient norms in higher layers during the back-propagation, which may lead to effective training. Exploiting the new findings, we propose a method that can provide both high stability and effective training by a simple modification of Post-LN. We conduct experiments on a wide range of text generation tasks. The experimental results demonstrate that our method outperforms Pre-LN, and enables stable training regardless of the shallow or deep layer settings. Our code is publicly available at https://github.com/takase/b2t_connection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Mix-LN, which uses Post-LN in early layers and Pre-LN in later layers, gives more uniform layer gradients and better LLM pretraining and fine-tuning results than Pre-LN or Post-LN alone.

Pith tools