Pith. sign in

REVIEW 5 cited by

MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.11733 v2 pith:YEBWTL32 submitted 2025-05-16 cs.CL

classification cs.CL
keywords reasoningdiagnosticclinicalllmsmedcasereasoningaccuracydatasetmedical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be objectively defined by the final answer, medical diagnosis requires both the outcome and the reasoning process to be accurate. Currently, widely used medical benchmarks like MedQA and MMLU assess only accuracy in the final answer, overlooking the quality and faithfulness of the clinical reasoning process. To address this limitation, we introduce MedCaseReasoning, the first open-access dataset for evaluating LLMs on their ability to align with clinician-authored diagnostic reasoning. The dataset includes 14,489 diagnostic question-and-answer cases, each paired with detailed reasoning statements derived from open-access medical case reports. We evaluate state-of-the-art reasoning LLMs on MedCaseReasoning and find significant shortcomings in their diagnoses and reasoning: for instance, the top-performing open-source model, DeepSeek-R1, achieves only 48% 10-shot diagnostic accuracy and mentions only 64% of the clinician reasoning statements (recall). However, we demonstrate that fine-tuning LLMs on the reasoning traces derived from MedCaseReasoning significantly improves diagnostic accuracy and clinical reasoning recall by an average relative gain of 29% and 41%, respectively. The open-source dataset, code, and models are available at https://github.com/kevinwu23/Stanford-MedCaseReasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Turning clinical guidelines into executable skill functions, refined with labeled cases, improves LLM diagnostic accuracy across four benchmarks and four backbones.

  2. MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Mid-stream alignment on next-step clinical decisions (MedUPS) improves LLM next-step accuracy on uncommon cases, but the size of the gain depends substantially on which LLM judge does the scoring.

  3. Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new 1,089-case multi-turn multimodal benchmark shows that even top AI medical diagnosticians are fully correct only ~34% of the time and frequently hallucinate reasoning.

  4. Auditing Evidence Use in Medical LLM Diagnosis

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Behavioral auditing of five medical LLMs shows most mined evidence interactions are clinically plausible, while adjudicated shortcut-like failures concentrate in negated or absent findings and clinically local evidence.

  5. Automatically Evolving Prompt Guidelines for Task-Specific Optimization

    cs.CL 2026-05 conditional novelty 6.0 of 10

    AGOPS automatically evolves task-specific prompt guidelines from reference answers and reports recovering 15.5–81.7% of the performance lost to underspecified prompts.

Pith tools