Pith. sign in

REVIEW 2 cited by

What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.07825 v1 pith:4UJ4DGFI submitted 2025-04-10 cs.CL

classification cs.CL
keywords common-sensereasoninghellaswagissueslanguagemodelsvaliditybenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Common-sense reasoning is a key language model capability because it encapsulates not just specific factual knowledge but rather general language and world understanding. Measuring common-sense reasoning, therefore, is crucial for language models of different sizes and applications. One of the most widely used benchmarks for evaluating such capabilities is HellaSwag; however, in this paper, we show that it has severe construct validity issues. These issues range from basic ungrammaticality and numerous typos to misleading prompts or equally correct options. Furthermore, we show that if models are evaluated only on answer texts, or with "Lorem ipsum dolor..." instead of the question, more than 65% of model predictions remain the same, and this cannot be attributed merely to contamination. Since benchmark scores are an essential part of model selection in both research and commercial applications, these validity issues can have severe consequences. In particular, knowing that taking benchmark scores at face value is ubiquitous, inadequate evaluation leads to ill-informed decisions about models. In this paper, we thoroughly investigate critical validity issues posed by HellaSwag and illustrate them with various evaluations using generative language models of different sizes. We argue that this benchmark does not accurately measure common-sense reasoning and, therefore, should not be used for evaluation in its current state. Based on the results of our study, we propose requirements that should be met by future common-sense reasoning benchmarks. In addition, we release GoldenSwag, a corrected subset of HellaSwag, which, to our belief, facilitates acceptable common-sense reasoning evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Commonsense benchmark scores are task-dependent predictors of downstream performance: they transfer reliably only to a few tasks (False Beliefs, TRIP), and revised benchmarks do not improve criterion validity.

  2. Human-AI Collaboration for Estimating Scientific Replicability

    cs.CY 2026-04 conditional novelty 6.0 of 10

    Hybrid human-AI prediction markets match or slightly outperform AI-only markets at forecasting scientific replication outcomes across six social science disciplines.

Pith tools