REVIEW 5 cited by
Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Pretrained contextual representation models (Peters et al., 2018; Devlin et al., 2018) have pushed forward the state-of-the-art on many NLP tasks. A new release of BERT (Devlin, 2018) includes a model simultaneously pretrained on 104 languages with impressive performance for zero-shot cross-lingual transfer on a natural language inference task. This paper explores the broader cross-lingual potential of mBERT (multilingual) as a zero shot language transfer model on 5 NLP tasks covering a total of 39 languages from various language families: NLI, document classification, NER, POS tagging, and dependency parsing. We compare mBERT with the best-published methods for zero-shot cross-lingual transfer and find mBERT competitive on each task. Additionally, we investigate the most effective strategy for utilizing mBERT in this manner, determine to what extent mBERT generalizes away from language specific features, and measure factors that influence cross-lingual transfer.
Forward citations
Cited by 5 Pith papers
-
Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks
Unicoder adds three cross-lingual pre-training tasks and a multi-language fine-tuning strategy to XLM, yielding modest gains (up to 0.7% in matched XNLI settings) and a new XQA benchmark.
-
Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation
A massively multilingual NMT encoder beats multilingual BERT in zero-shot cross-lingual transfer on 4 of 5 NLP tasks, but loses badly on named entity recognition.
-
Small and Practical BERT Models for Sequence Labeling
Distilling multilingual BERT into a 3-layer, 256-unit student yields a CPU-fast sequence labeler that is within about one F1 point of the teacher and beats a strong LSTM baseline.
-
Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
Using Layer 2 embeddings in Lugha-Llama, the paper reports a 28% relative increase in final-layer cosine similarity for Swahili-English pairs, but the supporting layer scan contains an internal contradiction and the c...
-
Investigating Multilingual NMT Representations at Scale
SVCCA analysis of a 103-language translation model shows encoder representations cluster by linguistic family, diverge by target language, and high-resource or related languages are more robust to fine-tuning.
Discussion (0). Continue with ORCID to comment.