REVIEW 5 cited by
MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be objectively defined by the final answer, medical diagnosis requires both the outcome and the reasoning process to be accurate. Currently, widely used medical benchmarks like MedQA and MMLU assess only accuracy in the final answer, overlooking the quality and faithfulness of the clinical reasoning process. To address this limitation, we introduce MedCaseReasoning, the first open-access dataset for evaluating LLMs on their ability to align with clinician-authored diagnostic reasoning. The dataset includes 14,489 diagnostic question-and-answer cases, each paired with detailed reasoning statements derived from open-access medical case reports. We evaluate state-of-the-art reasoning LLMs on MedCaseReasoning and find significant shortcomings in their diagnoses and reasoning: for instance, the top-performing open-source model, DeepSeek-R1, achieves only 48% 10-shot diagnostic accuracy and mentions only 64% of the clinician reasoning statements (recall). However, we demonstrate that fine-tuning LLMs on the reasoning traces derived from MedCaseReasoning significantly improves diagnostic accuracy and clinical reasoning recall by an average relative gain of 29% and 41%, respectively. The open-source dataset, code, and models are available at https://github.com/kevinwu23/Stanford-MedCaseReasoning.
Forward citations
Cited by 5 Pith papers
-
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
Turning clinical guidelines into executable skill functions, refined with labeled cases, improves LLM diagnostic accuracy across four benchmarks and four backbones.
-
MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models
Mid-stream alignment on next-step clinical decisions (MedUPS) improves LLM next-step accuracy on uncommon cases, but the size of the gain depends substantially on which LLM judge does the scoring.
-
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
A new 1,089-case multi-turn multimodal benchmark shows that even top AI medical diagnosticians are fully correct only ~34% of the time and frequently hallucinate reasoning.
-
Auditing Evidence Use in Medical LLM Diagnosis
Behavioral auditing of five medical LLMs shows most mined evidence interactions are clinically plausible, while adjudicated shortcut-like failures concentrate in negated or absent findings and clinically local evidence.
-
Automatically Evolving Prompt Guidelines for Task-Specific Optimization
AGOPS automatically evolves task-specific prompt guidelines from reference answers and reports recovering 15.5–81.7% of the performance lost to underspecified prompts.
Discussion (0). Continue with ORCID to comment.