Pith. sign in

REVIEW 7 cited by

Towards Multilingual LLM Evaluation for European Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08928 v2 pith:3YUWILYX submitted 2024-10-11 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesmultilingualbenchmarkseuropeanevaluationacrossllmslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and meaningful way across multiple European languages remains challenging, especially due to the scarcity of language-parallel multilingual benchmarks. We introduce a multilingual evaluation approach tailored for European languages. We employ translated versions of five widely-used benchmarks to assess the capabilities of 40 LLMs across 21 European languages. Our contributions include examining the effectiveness of translated benchmarks, assessing the impact of different translation services, and offering a multilingual evaluation framework for LLMs that includes newly created datasets: EU20-MMLU, EU20-HellaSwag, EU20-ARC, EU20-TruthfulQA, and EU20-GSM8K. The benchmarks and results are made publicly available to encourage further research in multilingual LLM evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Across 30 languages, commercial LLMs outscore all open-weight models in every EU language, and non-English service costs more and scores lower, suggesting equality requires resources beyond public web crawls.

  2. Meta-Learning Preferences for Multilingual LLM Alignment

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Meta-learning a shared initialization on multilingual preference data lets LLMs align to a new language from ~100 preference samples, with up to 28% win-rate gains over baselines.

  3. KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    KletterMix is a translated German corpus from English pretraining data that yields measurable gains on German downstream tasks in controlled pretraining experiments.

  4. Do LLMs exhibit the same commonsense capabilities across languages?

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs produce more commonsensical sentences in English than in Spanish, Dutch, or Valencian, across automatic, LLM-judge, and human evaluations on the new MULTICOM benchmark.

  5. Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety

    cs.CL 2025-06 conditional novelty 5.0 of 10

    NLP models trained on German accident reports outperform tabular-only models for accident classification, and LLM analysis suggests many fallback 'other' labels are parking accidents.

  6. From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A 2.7B German-first LLM trained cheaply on public data with language-specific quality filtering matches larger 7B models on German reasoning benchmarks and runs on-device.

  7. The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.

Pith tools