Pith. sign in

REVIEW 8 cited by

Multilingual Translation with Extensible Multilingual Pretraining and Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.00401 v1 pith:7LBRSBUB submitted 2020-08-02 cs.CL

classification cs.CL
keywords multilingualfinetuninglanguagesmodelstranslationpretrainedpretrainingscratch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work demonstrates the potential of multilingual pretraining of creating one model that can be used for various tasks in different languages. Previous work in multilingual pretraining has demonstrated that machine translation systems can be created by finetuning on bitext. In this work, we show that multilingual translation models can be created through multilingual finetuning. Instead of finetuning on one direction, a pretrained model is finetuned on many directions at the same time. Compared to multilingual models trained from scratch, starting from pretrained models incorporates the benefits of large quantities of unlabeled monolingual data, which is particularly important for low resource languages where bitext is not available. We demonstrate that pretrained models can be extended to incorporate additional languages without loss of performance. We double the number of languages in mBART to support multilingual machine translation models of 50 languages. Finally, we create the ML50 benchmark, covering low, mid, and high resource languages, to facilitate reproducible research by standardizing training and evaluation data. On ML50, we demonstrate that multilingual finetuning improves on average 1 BLEU over the strongest baselines (being either multilingual from scratch or bilingual finetuning) while improving 9.3 BLEU on average over bilingual baselines from scratch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 153 citations worldwide. Full citation record

  1. Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A new genderless-to-English benchmark shows that fine-tuning mBART-50 on carefully curated examples cuts gender stereotyping and pronoun-reasoning errors, beating larger proprietary systems on that benchmark.

  2. VTaMo: Video-Text Alignment Model for Sign Language Translation

    cs.CV 2026-07 accept novelty 6.5 of 10

    Explicit multi-granularity video–text alignment (OT + null token, orthogonal EMD, position contrastive) yields state-of-the-art gloss-free sign language translation on four benchmarks.

  3. CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    CC-Tuning fuses English feed-forward activations into non-English inputs during multilingual supervised fine-tuning, using a trainable Decision Maker and a least-squares Transform Matrix to simulate the connection at ...

  4. A Representation Level Analysis of NMT Model Robustness to Grammatical Errors

    cs.CL 2025-05 conditional novelty 6.0 of 10

    NMT encoders detect grammatical errors in early layers and move the error's representation toward the clean form in later layers; fine-tuning on noisy text increases reliance on the attention heads that do this work.

  5. A Large-Scale Benchmark for Vietnamese Sentence Paraphrases

    cs.CL 2025-02 conditional novelty 6.0 of 10

    ViSP is a 1.2M-pair Vietnamese sentence paraphrase benchmark, generated by Gemini and human-validated, with baseline and LLM evaluations showing BARTpho-word large and Meta-Llama-3.1-70B lead their respective categories.

  6. Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

    cs.CL 2026-08 conditional novelty 5.0 of 10

    Pruning multilingual NMT vocabularies to corpus-relevant tokens plus fine-tuning cuts memory by about 60% and matches or beats a dedicated English-Arabic model on COMET and TER.

  7. The first open machine translation system for the Chechen language

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A 171K-pair Chechen-Russian parallel corpus plus a fine-tuned NLLB-200 model are released, giving the first open Chechen-Russian translation system with human-evaluated quality near Google Translate.

  8. Two Spelling Normalization Approaches Based on Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Fine-tuned mT5 and mBART perform competitively on Spanish and Slovene historical spelling normalization, but character-based SMT achieves the best overall scores, confirming its continued state-of-the-art status.

Pith tools