Pith. sign in

REVIEW 15 cited by

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03304 v2 pith:5GDS2KFJ submitted 2024-12-04 cs.CL

classification cs.CL
keywords mmluevaluationquestionsbiasesculturalculturallyglobalmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from differences in language but also from the cultural knowledge required to interpret questions, reducing the practical utility of translated datasets like MMLU. Furthermore, translation often introduces artefacts that can distort the meaning or clarity of questions in the target language. A common practice in multilingual evaluation is to rely on machine-translated evaluation sets, but simply translating a dataset is insufficient to address these challenges. In this work, we trace the impact of both of these issues on multilingual evaluations and ensuing model performances. Our large-scale evaluation of state-of-the-art open and proprietary models illustrates that progress on MMLU depends heavily on learning Western-centric concepts, with 28% of all questions requiring culturally sensitive knowledge. Moreover, for questions requiring geographic knowledge, an astounding 84.9% focus on either North American or European regions. Rankings of model evaluations change depending on whether they are evaluated on the full portion or the subset of questions annotated as culturally sensitive, showing the distortion to model rankings when blindly relying on translated MMLU. We release Global MMLU, an improved MMLU with evaluation coverage across 42 languages -- with improved overall quality by engaging with compensated professional and community annotators to verify translation quality while also rigorously evaluating cultural biases present in the original dataset. This comprehensive Global MMLU set also includes designated subsets labeled as culturally sensitive and culturally agnostic to allow for more holistic, complete evaluation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Place Tokenizer Expansion for Pre-trained LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Continuing a model's own BPE merges and training only new embedding rows preserves quality while cutting token counts 2.4–4× for previously under-tokenized languages.

  2. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    cs.AI 2026-02 reject novelty 6.0 of 10

    Nearly half of 60 widely used LLM benchmarks show saturation, which increases with age; private test access does not prevent saturation, while expert curation does.

  3. A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Homophone normalization in Amharic training data can harm cross-lingual machine translation to Tigrinya and Ge'ez, but applying the normalization only at scoring time recovers BLEU gains.

  4. Mangosteen: An Open Thai Corpus for Language Model Pretraining

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An open 47B-token Thai pre-training corpus and a Thai-adapted data cleaning pipeline, with ablations showing quality gains and an 8B model that improves on Thai benchmarks.

  5. From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.

  6. NeoBabel: A Multilingual Open Tower for Visual Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.

  7. Evaluating Gemini in an arena for learning

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A blind expert-judged arena found Gemini 2.5 Pro preferred for tutoring in 73.2% of non-tied matchups against four competing models.

  8. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    cs.CL 2026-07 reject novelty 5.0 of 10

    Benchmark items are heterogeneous; the authors annotate them with three LLM judges and use the labels to build filtered subsets, but the validation evidence is weak and the comparison method is flawed.

  9. TASE: Token Awareness and Structured Evaluation for Multilingual Language Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    TASE benchmark shows LLMs lag humans on token-level and structural language tasks across Chinese, English, and Korean despite strong high-level performance.

  10. Against 'softmaxing' culture

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.

  11. Assessing the Role of Data Quality in Training Bilingual Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A quality filter trained only on English labels can select better French, German, and Chinese pretraining data, improving bilingual model performance and cutting the monolingual-bilingual gap to about 1%.

  12. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

  13. The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.

  14. LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multilingual visual question-answering benchmark across 11 languages and 5 social attributes, evaluated on 7 large multimodal models.

  15. Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network

    cs.CL 2026-07 unverdicted novelty 3.0 of 10

    A status report on a collaborative, student-run human-translation project to localise the MMLU benchmark into 11 European languages; no translated dataset or evaluation results are presented.

Pith tools