Pith. sign in

REVIEW 4 cited by

MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.10356 v2 pith:P3JRXDDI submitted 2025-04-14 cs.CL

classification cs.CL
keywords languagesmultilokolanguagemodelsbenchmarkfindllmsmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present MultiLoKo, a new benchmark for evaluating multilinguality in LLMs covering 31 languages. MultiLoKo consists of three partitions: a main partition consisting of 500 questions per language, separately sourced to be locally relevant to the specific language, and two translated partitions, containing human-authored translations from 30 non-English languages to English and vice versa. For comparison, we also release corresponding machine-authored translations. The data is equally distributed over two splits: a dev split and a blind, out-of-distribution test split. MultiLoKo can be used to study a variety of questions regarding the multilinguality of LLMs as well as meta-questions about multilingual benchmark creation. We compute MultiLoKo scores for 11 base and chat models marketed to be multilingual and study their average performance, their performance parity across languages, how much their ability to answer questions depends on the question language, and which languages are most difficult. None of the models we studied performs well on MultiLoKo, as indicated by low average scores as well as large differences between the best and worst scoring languages. Furthermore, we find a substantial effect of the question language, indicating sub-optimal knowledge transfer between languages. Lastly, we find that using local vs English-translated data can result in differences more than 20 points for the best performing models, drastically change the estimated difficulty of some languages. For using machines instead of human translations, we find a weaker effect on ordering of language difficulty, a larger difference in model rankings, and a substantial drop in estimated performance for all models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Across 30 languages, commercial LLMs outscore all open-weight models in every EU language, and non-English service costs more and scores lower, suggesting equality requires resources beyond public web crawls.

  2. MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

    cs.CL 2026-07 unverdicted novelty 7.0 of 10

    Multilingual LLMs show a reproducible 'Illusion of Cultural Alignment': they can be fluent in a language while lacking the culture's factual knowledge, and confidence, sampling, and retrieval do not fix it.

  3. MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A native-authored French, Spanish, and Chinese reasoning benchmark shows current LLMs score below 50% and improve by about 10% on math when questions are in English.

  4. Rethinking Cross-lingual Gaps from a Statistical Viewpoint

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Cross-lingual accuracy gaps in LLMs are dominated by higher response variance in target languages, not missing knowledge; ensembling and variance-reduction prompts shrink the gap.

Pith tools