Pith. sign in

REVIEW 2 cited by

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10647 v2 pith:EJP7CXB4 submitted 2025-03-02 cs.CL cs.AIcs.CYcs.HC

classification cs.CLcs.AIcs.CYcs.HC
keywords wereconsistencydiagnosticirrelevantchangesgeminiaddedclinical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and examination changes that preserved the diagnostic core. And the susceptibility was evaluated through embedding irrelevant but plausible narrative details while keeping the clinical evidence unchanged. For contextual awareness, patient history, lifestyle data, or diagnostic findings were added to shift the expected diagnosis. Physician reviewers then judged whether context-driven changes were clinically appropriate. Both models returned identical diagnoses across all equivalent variants and repeated queries (100% consistency). When irrelevant details were added, Gemini changed its diagnosis in 40.0% of cases and ChatGPT in 30.0%. ChatGPT responded to context more often than Gemini (77.8% vs. 55.6%), but a larger share of its changes were clinically inappropriate (33.3% vs. 22.2%). Gemini's context-driven changes were more often judged appropriate (66.7% vs. 55.6%). Consistency under controlled inputs did not protect either model from irrelevant manipulation or unjustified diagnostic shifts when context changed. Before LLMs can separate relevant from irrelevant input and flag insufficient evidence, their diagnostic use requires clinician oversight and structured safeguards.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

    cs.AI 2026-08 conditional novelty 5.0 of 10

    A three-step agentic workflow with LLM function calling and reflection improved glaucoma classification, CDR estimation, and repeatability over LLM-alone baselines, approaching specialist-level accuracy.

  2. Robust Fairness Vision-Language Learning for Medical Image Analysis

    cs.CV 2025-05 reject novelty 3.0 of 10

    A framework adding Dynamic Bad Pair Mining and Sinkhorn distance fairness loss to CLIP and BLIP-2 improves glaucoma diagnosis AUC on Harvard-FairVLMed, but fairness metrics worsen for several protected groups.

Pith tools