Pith. sign in

REVIEW 2 cited by

DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12566 v1 pith:GSIVB7ZT submitted 2020-10-23 cs.CL

classification cs.CL
keywords learningmultilingualrepresentationcross-lingualdict-mlmlanguageobjectivebetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained multilingual language models such as mBERT have shown immense gains for several natural language processing (NLP) tasks, especially in the zero-shot cross-lingual setting. Most, if not all, of these pre-trained models rely on the masked-language modeling (MLM) objective as the key language learning objective. The principle behind these approaches is that predicting the masked words with the help of the surrounding text helps learn potent contextualized representations. Despite the strong representation learning capability enabled by MLM, we demonstrate an inherent limitation of MLM for multilingual representation learning. In particular, by requiring the model to predict the language-specific token, the MLM objective disincentivizes learning a language-agnostic representation -- which is a key goal of multilingual pre-training. Therefore to encourage better cross-lingual representation learning we propose the DICT-MLM method. DICT-MLM works by incentivizing the model to be able to predict not just the original masked word, but potentially any of its cross-lingual synonyms as well. Our empirical analysis on multiple downstream tasks spanning 30+ languages, demonstrates the efficacy of the proposed approach and its ability to learn better multilingual representations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

    cs.CV 2026-08 conditional novelty 7.0 of 10

    Replacing an object's name in a caption with the image tokens of that object during pretraining gives explicit object-entity grounding, making MLLM alignment several times more data-efficient and boosting grounding an...

  2. Unsupervised Bilingual Lexicon Induction for Low Resource Languages

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Combining CSCBLI, linear transformation, and the iterative VecMap framework yields the top lexicon-induction accuracy on English with Sinhala, Tamil, and Punjabi, but the gains are small and the reported scores come f...

Pith tools