Pith. sign in

REVIEW 5 cited by

MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02257 v3 pith:RBAGG2DT submitted 2024-09-03 cs.CL cs.LG

classification cs.CLcs.LG
keywords mmlu-prollmsreasoningshortcutbiascorrectevaluationhigher-order
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing benchmarks for large language models (LLMs) increasingly struggle to differentiate between top-performing models, underscoring the need for more challenging evaluation frameworks. We introduce MMLU-Pro+, an enhanced benchmark building upon MMLU-Pro to assess shortcut learning and higher-order reasoning in LLMs. By incorporating questions with multiple correct answers across diverse domains, MMLU-Pro+ tests LLMs' ability to engage in complex reasoning and resist simplistic problem-solving strategies. Our results show that MMLU-Pro+ maintains MMLU-Pro's difficulty while providing a more rigorous test of model discrimination, particularly in multi-correct answer scenarios. We introduce novel metrics like shortcut selection ratio and correct pair identification ratio, offering deeper insights into model behavior and anchoring bias. Evaluations of six state-of-the-art LLMs reveal significant performance gaps, highlighting variations in reasoning abilities and bias susceptibility. We release the dataset and evaluation codes at \url{https://github.com/asgsaeid/mmlu-pro-plus}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Compute- and Data-Optimal Pretraining

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Pretraining loss obeys a single law in which repeated or paraphrased tokens count as η(N, data-per-parameter, expansion-ratio) fresh tokens, with total effective data saturating as derived tokens grow.

  2. Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A debate-based evaluation protocol on 50 MMLU-Pro questions: fine-tuning on the test set boosts standard accuracy from 50% to 82% but not debate win rates.

  3. Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Reliable deductive reasoning in AI requires replacing average-case statistical objectives with the exact learning criterion of universal correctness, a thesis supported by sample-complexity lower bounds showing statis...

  4. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

  5. Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks

    cs.CL 2025-07 reject novelty 3.0 of 10

    Across Arabic, English, and Kannada benchmarks, 4-bit and 8-bit quantization preserves most accuracy while aggressive pruning degrades larger multilingual models more than smaller ones.

Pith tools