Pith. sign in

REVIEW 1 cited by

Very Deep Transformers for Neural Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.07772 v2 pith:47VMV7KB submitted 2020-08-18 cs.CL

classification cs.CL
keywords bleumodelsdeeplayersmachineneuraltranslationvery
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore the application of very deep Transformer models for Neural Machine Translation (NMT). Using a simple yet effective initialization technique that stabilizes training, we show that it is feasible to build standard Transformer-based models with up to 60 encoder layers and 12 decoder layers. These deep models outperform their baseline 6-layer counterparts by as much as 2.5 BLEU, and achieve new state-of-the-art benchmark results on WMT14 English-French (43.8 BLEU and 46.4 BLEU with back-translation) and WMT14 English-German (30.1 BLEU).The code and trained models will be publicly available at: https://github.com/namisan/exdeep-nmt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Glinthawk: A Two-Tiered Architecture for Offline LLM Inference

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A two-tier inference system that offloads attention and KV cache to cheap CPU nodes raises offline LLM throughput about 6x and lowers hardware cost about 2.8x in a T4-based prototype.

Pith tools