Pith. sign in

REVIEW 3 cited by

Larger-Scale Transformers for Multilingual Masked Language Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.00572 v1 pith:AQQWSTSH submitted 2021-05-02 cs.CL

classification cs.CL
keywords modelslanguagelanguagesmodelxlm-raveragecross-linguallarger
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent work has demonstrated the effectiveness of cross-lingual language model pretraining for cross-lingual understanding. In this study, we present the results of two larger multilingual masked language models, with 3.5B and 10.7B parameters. Our two new models dubbed XLM-R XL and XLM-R XXL outperform XLM-R by 1.8% and 2.4% average accuracy on XNLI. Our model also outperforms the RoBERTa-Large model on several English tasks of the GLUE benchmark by 0.3% on average while handling 99 more languages. This suggests pretrained models with larger capacity may obtain both strong performance on high-resource languages while greatly improving low-resource languages. We make our code and models publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Language Transfer Capability of Decoder-only Architecture in Multilingual Neural Machine Translation

    cs.CL 2024-12 conditional novelty 7.0 of 10

    A two-stage decoder-only architecture with instruction-level contrastive learning improves zero-shot multilingual translation and closes most of the gap to encoder-decoder models.

  2. EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries

    cs.DB 2026-06 unverdicted novelty 6.0 of 10

    Query-driven table integration that uses Steiner-tree search to choose which joins LLMs must verify, reporting 30%+ accuracy gains at 5x lower LLM cost.

  3. When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A systematic comparison shows prompt-based LLMs underperform fine-tuned encoder models for reference-less translation quality estimation across eight low-resource language pairs.

Pith tools