Pith. sign in

REVIEW 7 cited by

Changing Answer Order Can Decrease MMLU Accuracy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19470 v2 pith:KGGC6CAI submitted 2024-06-27 cs.CL

classification cs.CL
keywords accuracymodelmodelsmmluanswerdecreasemultipleorder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. Most commonly, we use test accuracy averaged across multiple subtasks in order to rank models on leaderboards, to determine which model is best for our purposes. In this paper, we investigate the robustness of the accuracy measurement on a widely used multiple choice question answering dataset, MMLU. When shuffling the answer label contents, we find that all explored models decrease in accuracy on MMLU, but not every model is equally sensitive. These findings suggest a possible adjustment to the standard practice of leaderboard testing, where we additionally consider the percentage of examples each model answers correctly by random chance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Fine-tuning LLMs on an unseen language teaches syntax but fails to transfer semantic competence, leaving Python with up to a 19% performance advantage and no tested intervention closing the gap.

  2. RoToR: Towards More Reliable Responses for Order-Invariant Inputs

    cs.CL 2025-02 conditional novelty 6.0 of 10

    RoToR makes a frozen LLM order-invariant by circularly rotating a single global sort of segment position IDs, and Selective Routing combines it with the original model for mixed lists.

  3. Too Big to Fool: Resisting Deception in Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Larger language models resist misleading answer hints better than smaller ones while still following legitimate instructions.

  4. HARP: A challenging human-annotated math reasoning benchmark

    cs.LG 2024-12 conditional novelty 6.0 of 10

    HARP, a new benchmark of 5,409 US math competition problems with human solutions and choices, keeps frontier LLMs far from saturation: the best model scores 75.9% overall and only 41.1% on the hardest 197 problems.

  5. Benchmarking the Pedagogical Knowledge of Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors release an open benchmark of 1,143 pedagogical knowledge questions from Chilean teacher exams and report accuracy, cost, and size trade-offs for 97 large language models.

  6. MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

    cs.CL 2024-12 conditional novelty 5.0 of 10

    MMLU-CF is a new closed-test-set MCQ benchmark that uses rephrasing, choice shuffling, and 'None of the other choices' distractors to reduce data contamination, where GPT-4o scores 73.4% 5-shot.

  7. Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A 1.6B Arabic language model trained with synthetic multiple-choice instruction data outperforms 7B-13B models on several Arabic multiple-choice benchmarks.

Pith tools