Pith. sign in

REVIEW 7 cited by

On the Cross-lingual Transferability of Monolingual Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.11856 v3 pith:DQTCAILT submitted 2019-10-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords cross-lingualmultilinguallanguagelanguagesmodelsmonolingualabilityabstractions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art unsupervised multilingual models (e.g., multilingual BERT) have been shown to generalize in a zero-shot cross-lingual setting. This generalization ability has been attributed to the use of a shared subword vocabulary and joint training across multiple languages giving rise to deep multilingual abstractions. We evaluate this hypothesis by designing an alternative approach that transfers a monolingual model to new languages at the lexical level. More concretely, we first train a transformer-based masked language model on one language, and transfer it to a new language by learning a new embedding matrix with the same masked language modeling objective, freezing parameters of all other layers. This approach does not rely on a shared vocabulary or joint training. However, we show that it is competitive with multilingual BERT on standard cross-lingual classification benchmarks and on a new Cross-lingual Question Answering Dataset (XQuAD). Our results contradict common beliefs of the basis of the generalization ability of multilingual models and suggest that deep monolingual models learn some abstractions that generalize across languages. We also release XQuAD as a more comprehensive cross-lingual benchmark, which comprises 240 paragraphs and 1190 question-answer pairs from SQuAD v1.1 translated into ten languages by professional translators.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa

    cs.CL 2025-07 conditional novelty 7.0 of 10

    TyDi QA-WANA is a new 28,000-example QA benchmark covering 10 under-represented languages with long-context, information-seeking questions and baseline evaluations.

  2. skLEP: A Slovak General Language Understanding Benchmark

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A nine-task Slovak-language understanding benchmark with translated and newly curated datasets, plus the first broad fine-tuned model comparison for Slovak.

  3. Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    JQL trains small multilingual quality scorers from LLM judgments and human annotations, and filtering pretraining data with them improves downstream multilingual model performance over heuristic baselines.

  4. The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The authors manually annotated a 15,619-sentence Tagalog news treebank following Universal Dependencies and evaluated transformer-based dependency parsers against it, plus quality and topic analyses.

  5. Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs

    cs.CL 2026-02 reject novelty 5.0 of 10

    A three-stage SAE-plus-SVD steering recipe claims to make Hindi or Spanish the default language of an LLM at inference time, but the visible manuscript reports only expected, not measured, outcomes.

  6. Transfer of Structural Knowledge from Synthetic Languages

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A new synthetic language, flat_shuffle, transfers more structure to English fine-tuning than earlier synthetic bracket languages, though still far short of training on English from scratch.

  7. Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Selective pre-translation, translating only some prompt components into English, generally outperforms both full prompt translation and direct inference across tasks and languages, with the largest gains for low-resou...

Pith tools