Pith. sign in

REVIEW 4 cited by

Medical Large Language Model Benchmarks Should Prioritize Construct Validity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10694 v1 pith:5JRUKBBC submitted 2025-03-12 cs.CL

classification cs.CL
keywords benchmarksmedicalconstructvalidityevaluationclaimsclinicallanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation on competitive benchmarks; a tradition inherited from mainstream machine learning. But how do we separate real progress from a leaderboard flex? Medical LLM benchmarks, much like those in other fields, are arbitrarily constructed using medical licensing exam questions. For these benchmarks to truly measure progress, they must accurately capture the real-world tasks they aim to represent. In this position paper, we argue that medical LLM benchmarks should (and indeed can) be empirically evaluated for their construct validity. In the psychological testing literature, "construct validity" refers to the ability of a test to measure an underlying "construct", that is the actual conceptual target of evaluation. By drawing an analogy between LLM benchmarks and psychological tests, we explain how frameworks from this field can provide empirical foundations for validating benchmarks. To put these ideas into practice, we use real-world clinical data in proof-of-concept experiments to evaluate popular medical LLM benchmarks and report significant gaps in their construct validity. Finally, we outline a vision for a new ecosystem of medical LLM evaluation centered around the creation of valid benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Path to Multimodal Historical Reasoning: HistBench and HistAgent

    cs.AI 2025-05 conditional novelty 7.0 of 10

    HistAgent, a history-specialized agent, scores 27.54% pass@1 and 36.47% pass@2 on the new 414-question HistBench benchmark, surpassing generalist agents tested on the same data.

  2. Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Vision-language models frequently flip GBM versus metastasis diagnoses under evidence-preserving slice reordering and label-order swaps, with up to 67.8% flip rates, so static accuracy overstates clinical reliability.

  3. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0 of 10

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  4. Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks

    cs.HC 2025-07 conditional novelty 6.0 of 10

    Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...

Pith tools