Pith. sign in

REVIEW 3 cited by

Superhuman performance of a large language model on the reasoning tasks of a physician

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.10849 v3 pith:5426EBKH submitted 2024-12-14 cs.AI cs.CL

classification cs.AIcs.CL
keywords reasoningdiagnosticclinicalphysicianemergencyevaluationmedicalroom
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A global log for medical AI

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.

  2. Teaching large language models to reason like expert diagnosticians

    cs.AI 2025-09 conditional novelty 6.0 of 10

    An LLM agent and a 10-task benchmark built from 7,102 NEJM clinicopathologic cases push medical AI evaluation beyond final diagnosis accuracy.

  3. Towards medical AI misalignment: a preliminary study

    cs.CY 2025-05 conditional novelty 5.0 of 10

    A custom role-playing prompt called the Goofy Game made four major LLMs produce plausible but incorrect medical recommendations.

Pith tools