{"id":"ca17a827-dbcb-4f2c-b021-d0ca2417d227","arxiv_id":"2607.28677","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Passing medical exams does not make LLMs safe for autonomous triage: they fail to seek missing red flags, and current benchmarks do not test that.","lead":"This paper argues that large language models are not yet safe for autonomous medical triage, because they are trained to continue probable text rather than to hunt for missing red flags under incomplete information. It proposes that clinical safety claims must be backed by stress tests that withhold information and weigh a missed dangerous diagnosis more heavily than false alarms.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the evidence-gap claim is well supported; the mechanistic explanation is interesting but not load-bearing.","rationale":"The paper's central claim is an evidence-gap claim—'the evidence of safety does not yet exist' for autonomous LLM triage. This is a negative existential claim, best supported by systematic reviews (refs 15, 16) showing that only a tiny fraction of studies use real-world data and prospective designs, and by the paper's analysis that the few prospective studies that exist exclude the undifferentiated, high-acuity, incompletely disclosing patient. The reader's weakest_assumption focuses on the mechanistic premise that LLM training objectives causally explain the failure to seek missing information. I contend this mechanism is not load-bearing for the central claim. Even if the mechanism were false—i.e., if a well-designed prompting strategy could make current LLMs reliably seek missing red flags—the evidence-gap claim would still hold because no such system has yet demonstrated safe autonomous triage under realistic conditions. The paper itself acknowledges that structured elicitation can improve performance (ref 24) but argues that this is a layered policy, not a property of the objective, and that such policies have not been rigorously evaluated. Thus, the central claim survives even if the mechanism is wrong. The only way the central claim could fail is if a qualifying study demonstrating safety already exists but was missed. This is an empirical possibility, but the paper cites systematic reviews suggesting this is improbable. I therefore find no load-bearing concern that would change the ACCEPT verdict. Agreement is partial because the reader correctly identifies a soft spot in the paper's explanatory framing, but that soft spot is not load-bearing for the central claim.","tokens_in":9864,"tokens_out":10720,"duration_ms":102328,"concrete_test":"Conduct a PRISMA systematic review of all prospective studies of LLM-based triage systems operating without clinician oversight on undifferentiated patients with incomplete histories, applying the paper's inclusion/exclusion criteria. Operationalize safety via a pre-specified threshold for under-triage of must-not-miss diagnoses. If any study meets criteria and shows acceptable safety, the central evidence-gap claim would need revision; otherwise it stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an evidence-gap claim: no study has yet demonstrated that an LLM can safely perform autonomous triage of undifferentiated patients with incomplete histories. This claim is supported by the paper's systematic review of existing evaluations, showing that they fail to test this task due to pre-screening, clinician backstops, complete histories, or confidence gating. The reader's weakest assumption—that the training objective causally prevents reasoning from silence—is not load-bearing. Even if a prompting strategy or behavioral policy could fully mitigate the deficit, the evidence-gap claim would still hold: no such system has been demonstrated to be safe in a valid evaluation. The paper explicitly notes that structured elicitation policies exist (ref. 24) but argues they are not a property of the training objective and require proper evaluation. Thus, the central claim does not depend on the mechanism being true. The only potential soft spot is the completeness of the literature survey: if a qualifying study demonstrating safe autonomous triage exists, the claim would be false. The paper cites two systematic reviews showing that very few studies use real-world data and prospective designs, making this unlikely. No load-bearing concern identified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Perspective argues that no current evidence demonstrates that LLMs can safely perform autonomous triage of self-presenting, undifferentiated patients without a clinician in the loop. The paper distinguishes knowledge-oriented benchmarks (licensing exams, curated vignettes) from the real task of sequential decision-making under incomplete histories and asymmetric error costs, and argues that existing evaluations—including those of AMIE and autonomous primary-care systems—pre-filter patients, supply complete histories, or confidence-gate outputs, so they cannot detect failures of 'reasoning from silence.' It also identifies a cluster of assistant-like behaviors (credulity, agreeableness, miscalibration) that compound the deficit, and proposes three evaluation requirements: withhold information by design, score harm-weighted behaviors rather than top-k accuracy, and validate synthetic simulators through face, distributional, and predictive validity.","tokens_in":10096,"tokens_out":7872,"duration_ms":74956,"significance":"The evidence-gap claim is well supported: the paper marshals a randomized study showing a drop from 94.9% to <34.5% condition identification when real users provide histories [2], and two systematic reviews showing that only a small fraction of LLM evaluations use real-world prospective data [15,16]. If accepted, the paper provides a timely caution against premature deployment and a concrete, operationalizable evaluation agenda. Although it is a perspective rather than a new empirical study, its strengths include a falsifiable central claim, explicit positive proposals, and a clear breakdown of how specific design choices (pre-screening, confidence gating, complete histories) hide tail risks. The mechanistic explanation of the deficit is the least supported part, but it is not required for the central claim.","major_comments":[{"comment":"The paper presents a causal claim as established fact: 'a model trained to continue the most probable text has no native representation of the danger it was never told about' (§2), and 'the deeper problem is what the substrate was never built to do' (§1). The cited evidence ([11], [12]) demonstrates that current models behave this way, but it does not show that the training objective itself prevents mitigation; §5 acknowledges that structured elicitation policies exist (ref. 24). Since the central evidence-gap claim does not depend on this mechanistic explanation, please frame it as a hypothesis ('we hypothesize that...') rather than as a settled property of LLMs.","section":"Sections 1 and 2"},{"comment":"'All authors declare no conflict of interest' appears inconsistent with the author affiliations listing Atman Labs, a commercial AI company (author block, lines 3-4). Employment, equity, or funding from Atman Labs is a financial interest highly relevant to a Perspective arguing for stricter evidence before autonomous triage deployment. This must be disclosed or explicitly rebutted before publication.","section":"Declarations"}],"minor_comments":[{"comment":"The sentence reporting '94.9% ... but fewer than 34.5%' could clarify that both percentages refer to condition-identification rates, and that the first is on curated inputs while the second is on real-user dialogue; the current phrasing is slightly ambiguous.","section":"Section 1"},{"comment":"Reference [11] is a medRxiv preprint; the 1,000-consultation stress-test and the 97.5% accuracy figure rest on it. Please mark it as a preprint and temper 'a large stress-test confirms' to reflect the provisional status.","section":"Section 2"},{"comment":"'This journal's editors have put it more plainly still' is ambiguous because the manuscript does not identify the target journal; specify that ref. [29] is a Nature Medicine editorial.","section":"Section 5"},{"comment":"'Even the most advanced dynamic simulators still furnish the clinical facts' is unclear; specify that the simulator provides the clinical facts to the model rather than requiring the model to elicit them.","section":"Section 5"},{"comment":"The data availability statement claims a dataset 'used for this review' that will be shared; as this is a perspective with no original dataset, remove or clarify this statement.","section":"Declarations"},{"comment":"Consider defining 'autonomous triage' explicitly (degree of clinician backstop, setting, and population) at first use, since the paper discusses consumer symptom checkers, autonomous primary care systems, and ED agents, which differ in autonomy and risk.","section":"Section 1"}],"recommendation":"minor_revision","confidential_remarks":"I support publication after minor revision. The COI disclosure is the most urgent item; the mechanistic overstatement is easily corrected by reframing. The paper is an appropriate Perspective for a general medical or AI journal. I would not require additional original data, as the evidence-gap argument is built on a sufficient synthesis of published studies. The reliance on preprints for some behavioral claims should be flagged in the text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this one is that it's a perspective piece that actually makes a claim: no one has yet shown that an LLM can safely triage self-presenting patients with incomplete histories, and the evidence routinely cited for readiness doesn't test that. That claim is well supported. The paper isn't another 'LLMs are risky' editorial; it gives you a concrete mechanism (reasoning from silence) and a concrete evaluation program (withhold information, score the right objective, validate the simulator).\n\nWhat's genuinely new is the framing. The differential between 94.9% on curated inputs and <34.5% on real patient dialogue makes the point that the gap is in history-taking, not knowledge. The three-part simulator validation (face, distributional, predictive validity) is a practical contribution — it tells regulators what 'synthetic patient' evidence would have to look like. The paper is also honest: it flags the limitations of the studies it critiques, and it explicitly notes that structured elicitation policies exist and can help; they just aren't what the training objective provides.\n\nThe soft spots are real but not load-bearing. The claim that next-token prediction 'inherently' prevents models from reasoning about absence is asserted, not proven. References [1] and [11] show that with complete histories models perform well, but that doesn't show the training objective causes the failure; it might just be that current models haven't been optimized to ask, which is a different statement. The paper would be stronger if it said 'current LLMs do not reliably do this' rather than 'the substrate was never built to do it.' But the evidence-gap claim survives that softening — no one has yet demonstrated safe autonomous triage in a prospective real-world study, and the systematic reviews they cite support that.\n\nA minor quibble: the paper leans heavily on a single randomized study [2] for its headline numbers; that study looks well done, but the argument would benefit from acknowledging that one study isn't the whole evidence base. Also, the 'five behaviors' list is plausible but not empirically derived.\n\nOverall, this is a serious, well-reasoned piece. The citation pattern is appropriate; the one self-citation (Hillis & Payne) is peripheral. It deserves a real referee. I'd accept it as a perspective with light review, and I'd ask the authors to soften the mechanistic overreach in §1 and §2.","headline":"A well-argued evidence-gap perspective with a useful new frame; the causal training-objective claim is the soft underbelly, but the safety case doesn't rest on it.","tokens_in":10610,"tokens_out":3248,"would_cite":true,"duration_ms":27597,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No current large language model has demonstrated safe autonomous triage of undifferentiated patients.","keywords":["large language models","clinical decision support","autonomous triage","patient safety","reasoning from silence","benchmark validity","asymmetric cost","medical reasoning"],"falsifier":"A prospective randomized trial in a real emergency or primary-care setting where an LLM triages undifferentiated, incompletely disclosing patients, with the model required to ask for missing details and escalate high-harm cases; if the model's under-triage rate of must-not-miss diagnoses is comparable to or better than clinicians' on the same encounters, the central 'evidence does not exist' claim would be overturned for that system. A cheaper check: measure whether a model's differential width widens and its confidence falls as a fixed history is systematically reduced from complete to 20% co","tokens_in":9774,"feed_emoji":"🩺","tokens_out":3443,"duration_ms":31759,"temperature":0.7,"pith_summary":"This Perspective argues that the evidence used to claim large language models are ready for autonomous triage measures the wrong thing. Passing medical exams and scoring high on curated cases shows knowledge, not the ability to act safely when a patient's story is incomplete. The paper's core claim is that models trained to continue the most probable text are not optimized to seek the missing red flag, broaden the differential, or escalate when a must-not-miss diagnosis remains unexcluded. Current benchmarks supply complete, pre-cleaned histories and therefore cannot see this failure. The paper concludes that autonomous triage should not be deployed until systems pass evaluations that withhold information, score asymmetric harm, and validate simulators against real outcomes.","feed_headline":"No LLM has yet proven safe autonomous triage","feed_subtitle":"Models trained to write the most likely answer don't hunt for the red flag a patient never mentioned.","key_machinery":"The central mechanism is 'reasoning from silence': making safe inferences from diagnostically informative missing information—recognizing that what a patient has not said may reflect incomplete elicitation rather than true absence of disease. The paper uses this concept to explain why next-token-prediction training fails triage: the model conditions on what is present, not on what is absent, so absence exerts no force on confidence. The other load-bearing piece is the proposed evaluation design: withhold information, score the right objective with asymmetric cost, and validate the simulator in three layers (face, distributional, predictive).","core_discovery":"The paper's central claim is that the safety gap for autonomous triage is not medical knowledge but an objective mismatch: an LLM optimized to predict the most probable next token and to be helpful has no native representation of the cost of missing a rare, dangerous diagnosis. Under incomplete patient histories, models fail at the behaviors safe triage requires—widening the differential, seeking the missing red flag, lowering the threshold for escalation, and deferring when information is insufficient. Because existing evaluations use complete, expertly curated cases and confidence-gated populations that exclude the ambiguous patient, they cannot detect this failure. The paper proposes that","pith_inferences":["If the objective-mismatch account is right, then scale alone will not close the gap; the field may need a different training signal that rewards information seeking and harm-weighted deferral, not just helpfulness.","The same reasoning-from-silence logic likely applies beyond triage—to any high-stakes LLM decision under incomplete inputs, where the unstated detail is the one that matters.","A testable extension: build a benchmark that systematically varies information completeness and measures whether a model's differential width and escalation threshold respond to missingness; the paper predicts they will not.","The paper leaves open whether structured elicitation policies layered on top of current models can compensate; that is testable and is not settled by the paper's argument."],"forward_implications":["If correct, exam-passing and curated-benchmark accuracy cannot be used as evidence of readiness for autonomous triage, no matter how high the scores.","Deploying LLM symptom checkers or autonomous primary-care systems without evidence from incomplete-information, harm-weighted evaluations risks confident, plausible answers that close the case before a must-not-miss diagnosis is excluded.","Assistant-like behaviors—credulity, agreeableness, miscalibration—compound the core deficit, so evaluations must test resistance to minimized and adversarial patient histories, not just accuracy.","The paper's proposed standard—withhold information, score asymmetric cost, validate simulators—offers a concrete precondition that any autonomous triage claim must meet before it could justify deployment."],"fun_headline_variants":["LLM safety gap: trained to predict, not to seek red flags","Autonomous triage unsafe: LLMs miss the improbable must-not-miss","LLMs fail safe triage: they don't hunt for the unasked red flag","Objective mismatch: LLMs predict likely, not safe, diagnoses","Safety gap isn't knowledge but missing unmentioned red flags"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument depends on the premise that the training objective—predicting the most probable next token and being helpfully agreeable—is what causes models to ignore missing red flags, so that no amount of prompting or wrapping can fully fix the deficit.","fun_headline_variants_meta":{"raw":{"variants":["LLM safety gap: trained to predict, not to seek red flags","Autonomous triage unsafe: LLMs miss the improbable must-not-miss","LLMs fail safe triage: they don't hunt for the unasked red flag","Objective mismatch: LLMs predict likely, not safe, diagnoses","Safety gap isn't knowledge but missing unmentioned red flags"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000889,"raw_usage":{"total_tokens":3710,"prompt_tokens":822,"completion_tokens":2888,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":2803}},"tokens_in":566,"tokens_out":2888,"duration_ms":18383,"temperature":1.0,"reasoning_tokens":2803,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:42:23.689596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective randomized trial in a real emergency or primary-care setting where an LLM triages undifferentiated, incompletely disclosing patients, with the model required to ask for missing details and escalate high-harm cases; if the model's under-triage rate of must-not-miss diagnoses is comparable to or better than clinicians' on the same encounters, the central 'evidence does not exist' claim would be overturned for that system. A cheaper check: measure whether a model's differential width widens and its confidence falls as a fixed history is systematically reduced from complete to 20% co","supporting_citations":[],"review_version":1}