Pith. sign in

REVIEW 2 cited by

The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03193 v1 pith:DMW2B7ZG submitted 2021-06-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationlow-resourcetranslationlanguagesmachinemultilingualbenchmarkbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARC-Encoder: learning compressed text representations for large language models

    cs.CL 2025-10 conditional novelty 6.0 of 10

    ARC-Encoder pools queries in an encoder's last attention layer to produce compressed continuous representations that a frozen decoder consumes as token embeddings.

  2. Building a Functional Machine Translation Corpus for Kpelle

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper introduces the first claimed public English-Kpelle parallel corpus and shows that fine-tuning NLLB on it yields BLEU up to 30 for Kpelle-to-English.

Pith tools