Pith. sign in

REVIEW 7 cited by

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.19799 v1 pith:JOQLOVT2 submitted 2024-11-29 cs.CL

classification cs.CL
keywords multilinguallanguagesllmslanguagemanyregionalbenchmarkenglish
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (\ie, multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts. Our novel resource, INCLUDE, is a comprehensive knowledge- and reasoning-centric benchmark across 44 written languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A new 7,044-question native Sinhala exam benchmark shows the best LLM at 67.65% accuracy, with large drops on culturally specific subjects.

  2. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Lossy speculative-decoding verification splits into truncation-based and collaborative methods; truncation-based methods underperform their matched baselines, and capping draft overshoot preserves quality.

  3. Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.

  4. Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A 7B open-weight translation model matches or outperforms far larger commercial systems across 28 languages in automatic and human evaluations.

  5. Assessing the Role of Data Quality in Training Bilingual Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.

  6. AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    AgentScope 1.0 packages the components needed to build, evaluate, and deploy LLM agent applications into one developer framework.

  7. From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.

Pith tools