REVIEW 3 cited by
IndicXNLI: Evaluating Multilingual Inference for Indian Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce IndicXNLI, an NLI dataset for 11 Indic languages. It has been created by high-quality machine translation of the original English XNLI dataset and our analysis attests to the quality of IndicXNLI. By finetuning different pre-trained LMs on this IndicXNLI, we analyze various cross-lingual transfer techniques with respect to the impact of the choice of language models, languages, multi-linguality, mix-language input, etc. These experiments provide us with useful insights into the behaviour of pre-trained models for a diverse set of languages.
Forward citations
Cited by 3 Pith papers
-
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.
-
Analysis of Indic Language Capabilities in LLMs
A desk-research review finds that LLM performance is strongest for Hindi, Bengali, Marathi, Telugu, and Tamil, and recommends prioritizing these five languages for safety benchmarks.
-
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
Layer-pruned MahaBERT-v2 and Google-Muril models roughly match full models on Marathi headline and paragraph classification but lose ground on document classification, and they do not always beat same-size scratch-tra...
Discussion (0). Continue with ORCID to comment.