Pith. sign in

REVIEW 14 cited by

XNLI: Evaluating Cross-lingual Sentence Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1809.05053 v1 pith:HFEA623E submitted 2018-09-13 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagedatacross-lingualevaluationsentenceunderstandingxnlibaselines
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art natural language processing systems rely on supervision in the form of annotated data to learn competent models. These models are generally trained on data in a single language (usually English), and cannot be directly used beyond that language. Since collecting data in every language is not realistic, there has been a growing interest in cross-lingual language understanding (XLU) and low-resource cross-language transfer. In this work, we construct an evaluation set for XLU by extending the development and test sets of the Multi-Genre Natural Language Inference Corpus (MultiNLI) to 15 languages, including low-resource languages such as Swahili and Urdu. We hope that our dataset, dubbed XNLI, will catalyze research in cross-lingual sentence understanding by providing an informative standard evaluation task. In addition, we provide several baselines for multilingual sentence understanding, including two based on machine translation systems, and two that use parallel data to train aligned multilingual bag-of-words and LSTM encoders. We find that XNLI represents a practical and challenging evaluation suite, and that directly translating the test data yields the best performance among available baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 152 citations worldwide. Full citation record

  1. ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

    cs.CL 2025-10 conditional novelty 7.0 of 10

    ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.

  2. FLEXITOKENS: Flexible Tokenization for Evolving Language Models

    cs.CL 2025-07 unverdicted novelty 7.0 of 10

    FLEXITOKENS replaces rigid subword tokenizers and fixed-compression auxiliary losses with a simplified boundary-prediction objective in byte-level models, yielding lower over-fragmentation and up to 10-point gains on ...

  3. skLEP: A Slovak General Language Understanding Benchmark

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A nine-task Slovak-language understanding benchmark with translated and newly curated datasets, plus the first broad fine-tuned model comparison for Slovak.

  4. Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit

    cs.CL 2025-06 conditional novelty 6.0 of 10

    OMP sparse coding of donor token embeddings, with coefficients transferred to the base embedding space, preserves LLM performance after tokenizer replacement better than published zero-shot baselines, though simple he...

  5. LIMO: Less is More for Reasoning

    cs.CL 2025-02 unverdicted novelty 6.0 of 10

    LIMO achieves 63.3% on AIME24 and 95.6% on MATH500 via supervised fine-tuning on roughly 1% of the data used by prior models, supporting the claim that minimal strategic examples suffice when pre-training has already ...

  6. The Falcon Series of Open Language Models

    cs.CL 2023-11 conditional novelty 6.0 of 10

    Falcon-180B is a 180B-parameter open decoder-only model trained on 3.5 trillion tokens that approaches PaLM-2-Large performance at lower cost and is released with dataset extracts.

  7. Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs

    cs.CL 2026-02 reject novelty 5.0 of 10

    A three-stage SAE-plus-SVD steering recipe claims to make Hindi or Spanish the default language of an LLM at inference time, but the visible manuscript reports only expected, not measured, outcomes.

  8. TASE: Token Awareness and Structured Evaluation for Multilingual Language Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    TASE benchmark shows LLMs lag humans on token-level and structural language tasks across Chinese, English, and Korean despite strong high-level performance.

  9. Assessing the Role of Data Quality in Training Bilingual Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.

  10. LAGO: Few-shot Crosslingual Embedding Inversion Attacks via Language Similarity-Aware Graph Optimization

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LAGO shows that constraining alignment matrices of linguistically similar languages to be close improves few-shot cross-lingual embedding inversion accuracy over independent per-language baselines.

  11. Language-Specific Sentiment Polarity Biases in Encoder and Large Language Model Classification of Product Reviews

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    LLMs show negative polarity bias in French and encoder models show positive bias in Japanese when classifying product review sentiment.

  12. Cross-lingual Few-shot Learning for Persian Sentiment Analysis with Incremental Adaptation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Combining few-shot fine-tuning with incremental learning and regularization lets XLM-R and mDeBERTa reach about 96% accuracy on Persian sentiment analysis across five domains.

  13. Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions

    cs.CL 2025-06 conditional novelty 3.0 of 10

    LLM robustness research is organized into adversarial robustness, out-of-distribution robustness, and evaluation, with an accompanying GitHub collection of papers.

  14. Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

    cs.CL 2025-06

Pith tools