Pith. sign in

REVIEW 6 cited by

MuRIL: Multilingual Representations for Indian Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.10730 v2 pith:VBEKGOOL submitted 2021-03-19 cs.CL

classification cs.CL
keywords languagesmultilingualmurilindiatransliterateddatalanguagetext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

India is a multilingual society with 1369 rationalized languages and dialects being spoken across the country (INDIA, 2011). Of these, the 22 scheduled languages have a staggering total of 1.17 billion speakers and 121 languages have more than 10,000 speakers (INDIA, 2011). India also has the second largest (and an ever growing) digital footprint (Statista, 2020). Despite this, today's state-of-the-art multilingual systems perform suboptimally on Indian (IN) languages. This can be explained by the fact that multilingual language models (LMs) are often trained on 100+ languages together, leading to a small representation of IN languages in their vocabulary and training data. Multilingual LMs are substantially less effective in resource-lean scenarios (Wu and Dredze, 2020; Lauscher et al., 2020), as limited data doesn't help capture the various nuances of a language. One also commonly observes IN language text transliterated to Latin or code-mixed with English, especially in informal settings (for example, on social media platforms) (Rijhwani et al., 2017). This phenomenon is not adequately handled by current state-of-the-art multilingual LMs. To address the aforementioned gaps, we propose MuRIL, a multilingual LM specifically built for IN languages. MuRIL is trained on significantly large amounts of IN text corpora only. We explicitly augment monolingual text corpora with both translated and transliterated document pairs, that serve as supervised cross-lingual signals in training. MuRIL significantly outperforms multilingual BERT (mBERT) on all tasks in the challenging cross-lingual XTREME benchmark (Hu et al., 2020). We also present results on transliterated (native to Latin script) test sets of the chosen datasets and demonstrate the efficacy of MuRIL in handling transliterated data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

    cs.CL 2026-08 conditional novelty 6.0 of 10

    For Hindi vocabulary extension of a 30B LLM, the best embedding initialization is uniform subword averaging with Hindi norm calibration on the input and character-length-weighted averaging on the output, cutting conti...

  2. MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new human-corrected Marathi paraphrase detection corpus with 8,000 pairs in five difficulty buckets, benchmarked with BERT models, with MahaBERT reaching 88.7% F1.

  3. Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

    cs.CL 2026-07 conditional novelty 5.0 of 10

    A gated fusion head that conditions English toxicity, Indic abuse, and rule-based severity scores on the text context improves code-mixed abuse detection in 10/12 in-domain and 7/8 transfer comparisons.

  4. SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The authors release sense-annotated WSD/WiC datasets for ten low-resource languages and report that English-based zero-shot transfer often beats small in-language fine-tuning, while mixed training usually helps.

  5. Enhancing Hindi NER in Low Context: A Comparative study of Transformer-based models with vs. without Retrieval Augmentation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Adding Wikipedia context to Hindi NER inputs lifts XLM-R from 0.50 to 0.72 macro F1, but helps MuRIL only slightly and hurts or leaves unchanged all Llama-based systems.

  6. Cyberbullying Detection in Hinglish Text Using MURIL and Explainable AI

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A MURIL-based classifier outperforms RoBERTa, IndicBERT, and several published baselines on six Hinglish cyberbullying datasets, with reported gains of 1.36 to 13.07 percentage points.

Pith tools