Pith. sign in

REVIEW 6 cited by

Building Machine Translation Systems for the Next Thousand Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.03983 v3 pith:74SKHVPQ submitted 2022-05-09 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesmodelsbuildingsystemsdatasetsdevelopingleveragingmachine
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper we share findings from our effort to build practical machine translation (MT) systems capable of translating across over one thousand languages. We describe results in three research domains: (i) Building clean, web-mined datasets for 1500+ languages by leveraging semi-supervised pre-training for language identification and developing data-driven filtering techniques; (ii) Developing practical MT models for under-served languages by leveraging massively multilingual models trained with supervised parallel data for over 100 high-resource languages and monolingual datasets for an additional 1000+ languages; and (iii) Studying the limitations of evaluation metrics for these languages and conducting qualitative analysis of the outputs from our MT models, highlighting several frequent error modes of these types of models. We hope that our work provides useful insights to practitioners working towards building MT systems for currently understudied languages, and highlights research directions that can complement the weaknesses of massively multilingual models in data-sparse settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Per-step retrieval of solved exemplars injected into the reasoning trace improves test-time scaling accuracy, with up to 13.4 absolute points gained on AIME 2025.

  2. TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    TopXGen generates topic-diverse synthetic parallel data by prompting an LLM to write in low-resource languages and backtranslating to English, improving MT in ICL and fine-tuning across ten languages.

  3. FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An adaptive, per-language data filtering and deduplication pipeline produces multilingual LLM pre-training corpora that beat prior public datasets on 11 of 14 evaluated languages, and a 20TB, 1,868 language-script dat...

  4. Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Human-translated benchmarks in eight African languages show GPT-4o accuracy is 12 to 20 percentage points below English, and fine-tuning on translated data recovers part of the gap.

  5. PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks

    cs.CL 2024-12 conditional novelty 6.0 of 10

    PromptRefine uses alternating minimization over language-specific retrievers plus diversity-aware DPP fine-tuning to select cross-lingual in-context examples, improving few-shot generation in low-resource Indic languages.

  6. Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Back translation with dataset selection raises English-Luganda NMT BLEU scores by about 10 points on a new test set, though the comparison with previous benchmarks is not on a shared test set.

Pith tools