Pith. sign in

REVIEW 5 cited by

Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.09077 v2 pith:MR6TYT4P submitted 2019-04-19 cs.CL

classification cs.CL
keywords cross-lingualmbertlanguagetransferbertdevlinlanguagesmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pretrained contextual representation models (Peters et al., 2018; Devlin et al., 2018) have pushed forward the state-of-the-art on many NLP tasks. A new release of BERT (Devlin, 2018) includes a model simultaneously pretrained on 104 languages with impressive performance for zero-shot cross-lingual transfer on a natural language inference task. This paper explores the broader cross-lingual potential of mBERT (multilingual) as a zero shot language transfer model on 5 NLP tasks covering a total of 39 languages from various language families: NLI, document classification, NER, POS tagging, and dependency parsing. We compare mBERT with the best-published methods for zero-shot cross-lingual transfer and find mBERT competitive on each task. Additionally, we investigate the most effective strategy for utilizing mBERT in this manner, determine to what extent mBERT generalizes away from language specific features, and measure factors that influence cross-lingual transfer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks

    cs.CL 2019-09 conditional novelty 6.0 of 10

    Unicoder adds three cross-lingual pre-training tasks and a multi-language fine-tuning strategy to XLM, yielding modest gains (up to 0.7% in matched XNLI settings) and a new XQA benchmark.

  2. Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation

    cs.CL 2019-09 conditional novelty 6.0 of 10

    A massively multilingual NMT encoder beats multilingual BERT in zero-shot cross-lingual transfer on 4 of 5 NLP tasks, but loses badly on named entity recognition.

  3. Small and Practical BERT Models for Sequence Labeling

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Distilling multilingual BERT into a 3-layer, 256-unit student yields a CPU-fast sequence labeler that is within about one F1 point of the teacher and beats a strong LSTM baseline.

  4. Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning

    cs.CL 2025-06 reject novelty 5.0 of 10

    Using Layer 2 embeddings in Lugha-Llama, the paper reports a 28% relative increase in final-layer cosine similarity for Swahili-English pairs, but the supporting layer scan contains an internal contradiction and the c...

  5. Investigating Multilingual NMT Representations at Scale

    cs.CL 2019-09 conditional novelty 5.0 of 10

    SVCCA analysis of a 103-language translation model shows encoder representations cluster by linguistic family, diverge by target language, and high-resource or related languages are more robust to fine-tuning.

Pith tools