Pith. sign in

REVIEW 6 cited by

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.11080 v5 pith:B5H55RVH submitted 2020-03-24 cs.CL cs.LG

classification cs.CLcs.LG
keywords tasksbenchmarkmodelsacrosscross-linguallanguagesmultilingualbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders XTREME benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KatotohananQA is a Filipino translation of TruthfulQA; seven LLMs scored 94.72% in English versus 83.87% in Filipino, with GPT-5 and GPT-5 mini showing the smallest gap.

  2. When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Projecting away the estimated Modern Standard Arabic subspace during fine-tuning improves generation across 25 Arabic dialects by up to +4.9 chrF++, evidence that subspace dominance by a high-resource variety restrict...

  3. Beyond Literal Token Overlap: Token Alignability for Multilinguality

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new metric based on subword token alignment predicts cross-lingual transfer in multilingual models better than literal token overlap, especially for different-script language pairs.

  4. Find Central Dogma Again: Leveraging Multilingual Transfer in Large Language Models

    q-bio.GN 2025-02 reject novelty 5.0 of 10

    A GPT-2 model fine-tuned on multilingual sentence similarity achieves at best 81% accuracy on classifying matching vs non-matching DNA-protein pairs, but the result is highly seed-dependent and only with an easy test set.

  5. Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource Languages

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Federated averaging of prompt embeddings from a frozen multilingual model improves accuracy on some low-resource tasks (XNLI) but not consistently on others (MasakhaNEWS).

  6. Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

    cs.CL 2025-06

Pith tools