Pith. sign in

REVIEW 3 major objections 1 minor 22 references

Medical LLMs change clinical advice and can suggest harm when prompts are slightly reworded

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-27 21:39 UTC pith:3OVQ2TMK

load-bearing objection This paper benchmarks prompt sensitivity on MedMCQA for medical LLMs but the tested perturbations lack any check against real clinical query distributions. the 3 major comments →

arxiv 2606.07237 v1 pith:3OVQ2TMK submitted 2026-06-05 cs.CL cs.AIcs.LG

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

classification cs.CL cs.AIcs.LG
keywords large language modelshealthcareprompt sensitivityadversarial robustnessclinical reasoningMedMCQA
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tests how general and medical LLMs respond to small changes in how medical questions are phrased. It applies both everyday rewordings and deliberately tricky manipulations to models such as GPT-3.5, Llama3, ClinicalBERT and BioLlama3 on the MedMCQA dataset. Results show that minor phrasing shifts often produce different answers or dangerous recommendations like wrong drug doses. The models handle simple word swaps better than sentence reorderings or added misleading context. This pattern holds for both broad and domain-specific models, indicating that current systems lack the stability needed for clinical use.

Core claim

Medical LLMs are not intrinsically safe. Even minor variations in phrasing can alter clinical advice, and targeted adversarial prompts can provoke harmful outputs. Models tend to show resilience to simple lexical substitutions or paraphrasing, but they often break down under syntactic reordering or misleading contextual cues. This fragility is evident across both general-purpose and domain-specific LLMs, with adversarial manipulations leading to clinically dangerous outputs such as recommending incorrect dosages or omitting critical findings.

What carries the argument

Systematic sensitivity analysis that divides prompt changes into natural and adversarial categories and measures effects on consistency, accuracy, and reliability using the MedMCQA benchmark.

Load-bearing premise

The natural and adversarial prompt changes tested here match the variations that would actually occur when clinicians interact with these models.

What would settle it

Collect real prompts from clinicians using the models in practice and check whether the same inconsistencies and harmful outputs appear at similar rates.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Clinicians cannot rely on models that shift diagnoses or treatments based on reworded questions.
  • Adversarial inputs can produce outputs that recommend incorrect dosages or omit key medical findings.
  • The same sensitivity appears in both general-purpose models and those fine-tuned on medical data.
  • Unpredictable behavior makes these models unsuitable for high-stakes healthcare decisions without additional controls.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Standardizing prompt formats before model input could reduce some of the observed variability.
  • Testing models against logs of actual clinical conversations would provide a stronger check than benchmark perturbations alone.
  • Adding output verification steps or ensemble methods might limit the impact of prompt-induced errors.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The paper claims that both general-purpose (GPT-3.5, Llama3) and medical-specific LLMs (ClinicalBERT, BioLlama3, BioBERT) are highly sensitive to lexical, syntactic, natural, and adversarial prompt perturbations on the MedMCQA benchmark, resulting in inconsistent accuracy, altered clinical reasoning, and potentially harmful outputs such as incorrect dosages; it concludes that medical LLMs are not intrinsically safe for healthcare tasks including clinical QA, diagnosis support, and report summarization.

Significance. If the empirical results hold with proper quantification and validation, the work would provide concrete evidence of prompt fragility in safety-critical domains, strengthening calls for robustness benchmarks and potentially affecting deployment guidelines for clinical LLMs. The direct measurement approach on a standard benchmark is a positive aspect of the design.

major comments (3)
  1. [Abstract] Abstract: the findings are asserted without any quantitative results, accuracy deltas, statistical tests, number of examples tested, or details on perturbation generation procedures, preventing verification that the observed sensitivity supports the central claim of non-intrinsic safety.
  2. [Methods] Methods/Experimental Setup: the natural and adversarial perturbations are categorized but no validation is provided that they are representative of actual clinical prompt variations arising in EHR systems, diagnostic workflows, or clinician-LLM interactions; MCQA accuracy shifts are not shown to equate to changed clinical advice in the open-ended tasks referenced in the abstract.
  3. [Results] Results/Discussion: the evaluation is confined to MedMCQA (a multiple-choice QA dataset), yet claims about altered diagnoses, omitted findings, and harmful outputs in diagnosis support and summarization tasks lack corresponding experiments or justification for generalizing from closed-ended MCQA performance.
minor comments (1)
  1. Include at least one table or figure summarizing perturbation categories with concrete examples to improve clarity of the experimental design.

Simulated Author's Rebuttal

3 responses · 0 unresolved

Thank you for the constructive feedback. We address each major comment below and indicate revisions to improve quantification, scope clarification, and justification.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the findings are asserted without any quantitative results, accuracy deltas, statistical tests, number of examples tested, or details on perturbation generation procedures, preventing verification that the observed sensitivity supports the central claim of non-intrinsic safety.

    Authors: We agree the abstract is currently qualitative. The revision will add quantitative elements: accuracy deltas across models, statistical test results, number of examples tested (from MedMCQA), and perturbation generation details to substantiate the sensitivity claims. revision: yes

  2. Referee: [Methods] Methods/Experimental Setup: the natural and adversarial perturbations are categorized but no validation is provided that they are representative of actual clinical prompt variations arising in EHR systems, diagnostic workflows, or clinician-LLM interactions; MCQA accuracy shifts are not shown to equate to changed clinical advice in the open-ended tasks referenced in the abstract.

    Authors: Perturbations draw from documented clinical language variations and prompt engineering literature. We will add citations and a limitations discussion on representativeness. We will clarify MCQA as a proxy for reasoning consistency while noting it does not directly equate to open-ended advice. revision: partial

  3. Referee: [Results] Results/Discussion: the evaluation is confined to MedMCQA (a multiple-choice QA dataset), yet claims about altered diagnoses, omitted findings, and harmful outputs in diagnosis support and summarization tasks lack corresponding experiments or justification for generalizing from closed-ended MCQA performance.

    Authors: We will revise claims to center on clinical QA tasks and add justification for implications to related tasks based on observed reasoning changes, while toning down unsupported generalizations about diagnosis support and summarization. revision: yes

Circularity Check

0 steps flagged

No circularity: empirical benchmarking with direct measurements

full rationale

The paper is a straightforward empirical evaluation of LLM robustness on MedMCQA under categorized prompt perturbations. No derivation chain, fitted parameters renamed as predictions, self-definitional constructs, or load-bearing self-citations appear in the provided abstract or described methodology. Central claims rest on observed accuracy/consistency shifts from explicit experiments rather than any reduction to prior inputs or ansatzes. This matches the default expectation for non-circular empirical studies.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The claims depend on the representativeness of the benchmark and the perturbation categories for real-world clinical scenarios.

axioms (1)
  • domain assumption The MedMCQA benchmark adequately represents clinical reasoning tasks in healthcare.
    The study relies on this benchmark for all evaluations without additional validation mentioned.

pith-pipeline@v0.9.1-grok · 5773 in / 1047 out tokens · 26804 ms · 2026-06-27T21:39:42.305053+00:00 · methodology

0 comments
read the original abstract

Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization. Despite their promise, these models remain highly sensitive to subtle prompt perturbations, both lexical and syntactic, posing serious risks in safety-critical clinical applications. In this study, we conduct a systematic sensitivity analysis to evaluate the robustness of both general-purpose (e.g., GPT-3.5, Llama3) and medical-specific LLMs (e.g., ClinicalBERT, BioLlama3, BioBERT) using the MedMCQA benchmark. We categorize perturbations into natural and adversarial types and examine their effect on model consistency, accuracy, and reliability in clinical reasoning tasks. Our findings reveal that medical LLMs are not intrinsically safe. Even minor variations in phrasing can alter clinical advice, and targeted adversarial prompts can provoke harmful outputs. In high-stakes settings like healthcare, such unpredictability is unacceptable-models that change diagnoses due to reworded inputs or hallucinate medications when slightly rephrased cannot be reliably trusted by clinicians. While models tend to show resilience to simple lexical substitutions or paraphrasing, they often break down under syntactic reordering or misleading contextual cues. This fragility is evident across both general-purpose and domain-specific LLMs. Notably, adversarial manipulations can lead to clinically dangerous outputs, such as recommending incorrect dosages or omitting critical findings.

Figures

Figures reproduced from arXiv: 2606.07237 by Mahdi Alkaeed.

Figure 1
Figure 1. Figure 1: This methodological pipeline outlines a comprehensive robustness evaluation framework for assessing the sensitivity and stability [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of accurate and misinformation-attacked medical advice for acute limb ischemia. Lexical and syntactic alterations [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Performance of models across perturbation types. Heatmap shows scores for all metrics for each model and study, with ‘N’ [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Average similarity scores on 100 MedMCQA prompts across increasing lexical perturbation levels, showing that BioLLaMA2-7B [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 5 canonical work pages

  1. [1]

    Zhang, M

    P. Zhang, M. N. Kamel Boulos, Generative ai in medicine and healthcare: Promises, opportunities and challenges, Future Internet 15 (9) (2023) 286

  2. [2]

    arXiv preprint arXiv:2310.10844 (2023)

    E. Shayegani, M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, N. Abu-Ghazaleh, Survey of vulnerabilities in large language models revealed by adversarial attacks, arXiv preprint arXiv:2310.10844 (2023)

  3. [3]

    F. M. Polo, R. Xu, L. Weber, M. Silva, O. Bhardwaj, L. Choshen, A. F. de Oliveira, Y. Sun, M. Yurochkin, Efficient multi-prompt evaluation of llms, Advances in Neural Information Processing Systems 37 (2024) 22483–22512

  4. [4]

    S. Qi, Z. Cao, J. Rao, L. Wang, J. Xiao, X. Wang, What is the limitation of multimodal llms? a deeper look into multimodal llms through prompt probing, Information Processing & Management 60 (6) (2023) 103510

  5. [5]

    B. Guan, T. Roosta, P. Passban, M. Rezagholizadeh, The order effect: Investigating prompt sensitivity in closed-source llms, arXiv preprint arXiv:2502.04134 (2025)

  6. [6]

    W. J. Bolton, R. Poyiadzi, E. R. Morrell, G. v. B. G. Bueno, L. Goetz, Rambla: A framework for evaluating the reliability of llms as assistants in the biomedical domain, arXiv preprint arXiv:2403.14578 (2024)

  7. [7]

    R. O. Ness, K. Matton, H. Helm, S. Zhang, J. Bajwa, C. E. Priebe, E. Horvitz, Medfuzz: Exploring the robustness of large language models in medical question answering, arXiv preprint arXiv:2406.06573 (2024)

  8. [8]

    T. Han, S. Nebelung, F. Khader, T. Wang, G. Müller-Franzes, C. Kuhl, S. Försch, J. Kleesiek, C. Haarburger, K. K. Bressem, et al., Medical large language models are susceptible to targeted misinformation attacks, NPJ digital medicine 7 (1) (2024) 288. 11

  9. [9]

    J. Zhuo, S. Zhang, X. Fang, H. Duan, D. Lin, K. Chen, Prosa: Assessing and understanding the prompt sensitivity of llms, in: Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 1950–1976

  10. [10]

    Y. Wang, Y. Zhao, Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models, arXiv preprint arXiv:2406.11020 (2024)

  11. [11]

    Errica, D

    F. Errica, D. Sanvito, G. Siracusano, R. Bifulco, What did i do wrong? quantifying llms’ sensitivity and consistency to prompt engineering, in: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 1543–1558

  12. [12]

    Salinas, F

    A. Salinas, F. Morstatter, The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance, in: Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 4629–4651

  13. [13]

    B. Cao, D. Cai, Z. Zhang, Y. Zou, W. Lam, On the worst prompt performance of large language models, Advances in Neural Information Processing Systems 37 (2024) 69022–69042

  14. [14]

    Sclar, Y

    M. Sclar, Y. Choi, Y. Tsvetkov, A. Suhr, Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting, in: International Conference on Learning Representations, Vol. 2024, 2024, pp. 25055–25083

  15. [15]

    Pezeshkpour, E

    P. Pezeshkpour, E. Hruschka, Large language models sensitivity to the order of options in multiple-choice questions, in: Findings of the Association for Computational Linguistics: NAACL 2024, 2024, pp. 2006–2017

  16. [16]

    C. Yan, X. Fu, Y. Xiong, T. Wang, S. C. Hui, J. Wu, X. Liu, Llm sensitivity evaluation framework for clinical diagnosis, in: Proceedings of the 31st International Conference on Computational Linguistics, 2025, pp. 3083–3094

  17. [17]

    Moradi, M

    M. Moradi, M. Samwald, Improving the robustness and accuracy of biomedical language models through adversarial training, Journal of Biomedical Informatics 132 (2022) 104114

  18. [18]

    A. M. Ceballos-Arroyo, M. Munnangi, J. Sun, K. Zhang, J. Mcinerney, B. C. Wallace, S. Amir, Open (clinical) llms are sensitive to instruction phrasings, in: Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, 2024, pp. 50–71

  19. [19]

    Beede, E

    E. Beede, E. Baylor, F. Hersch, A. Iurchenko, L. Wilcox, P. Ruamviboonsuk, L. M. Vardoulakis, A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy, in: Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–12

  20. [20]

    C. Pais, J. Liu, R. Voigt, V. Gupta, E. Wade, M. Bayati, Large language models for preventing medication direction errors in online pharmacies, Nature Medicine (2024) 1–9

  21. [21]

    Zhang, L

    A. Zhang, L. Xing, J. Zou, J. C. Wu, Shifting machine learning for healthcare from development to deployment and from models to data, Nature Biomedical Engineering 6 (12) (2022) 1330–1345

  22. [22]

    P. Zhan, Z. Xu, Q. Tan, J. Song, R. Xie, Unveiling the lexical sensitivity of llms: Combinatorial optimization for prompt enhancement, in: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 5128–5154. 12