Pith. sign in

REVIEW 2 cited by

Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06644 v1 pith:DWGTHPAA submitted 2024-04-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagellmspersianchallengedataevaluationkhayyamcomprehension
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating Large Language Models (LLMs) is challenging due to their generative nature, necessitating precise evaluation methodologies. Additionally, non-English LLM evaluation lags behind English, resulting in the absence or weakness of LLMs for many languages. In response to this necessity, we introduce Khayyam Challenge (also known as PersianMMLU), a meticulously curated collection comprising 20,192 four-choice questions sourced from 38 diverse tasks extracted from Persian examinations, spanning a wide spectrum of subjects, complexities, and ages. The primary objective of the Khayyam Challenge is to facilitate the rigorous evaluation of LLMs that support the Persian language. Distinctive features of the Khayyam Challenge are (i) its comprehensive coverage of various topics, including literary comprehension, mathematics, sciences, logic, intelligence testing, etc., aimed at assessing different facets of LLMs such as language comprehension, reasoning, and information retrieval across various educational stages, from lower primary school to upper secondary school (ii) its inclusion of rich metadata such as human response rates, difficulty levels, and descriptive answers (iii) its utilization of new data to avoid data contamination issues prevalent in existing frameworks (iv) its use of original, non-translated data tailored for Persian speakers, ensuring the framework is free from translation challenges and errors while encompassing cultural nuances (v) its inherent scalability for future data updates and evaluations without requiring special human effort. Previous works lacked an evaluation framework that combined all of these features into a single comprehensive benchmark. Furthermore, we evaluate a wide range of existing LLMs that support the Persian language, with statistical analyses and interpretations of their outputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A new 7,044-question native Sinhala exam benchmark shows the best LLM at 67.65% accuracy, with large drops on culturally specific subjects.

  2. MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language

    cs.CL 2025-08 conditional novelty 6.0 of 10

    MELAC introduces 19 Persian and Iranian-culture evaluation datasets and benchmarks 41 LLMs, showing weak performance on Iranian-specific content.

Pith tools