{"id":"eb5ef41f-37b3-4e1f-bb3f-9738accd3ecf","arxiv_id":"2504.18919","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In a randomized trial, people using GPT-4o, Llama 3, or Command R+ for medical scenarios identified relevant conditions and dispositions no better than controls, despite the models performing much better alone.","lead":"A controlled trial with 1,298 people found that having an LLM chat assistant did not help people identify medical conditions or decide what care to seek, and in some ways hurt compared with their usual resources. This matters because medical benchmarks and simulated-patient tests did not predict these failures, so real human testing may be needed before such systems are deployed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Condition-identification scoring may systematically undercount LLM users' lay paraphrases, because fuzzy string matching against physician gold lists favors exact terms more likely in control internet-search responses; re-scoring with blinded semantic judgment is needed.","rationale":"The reader identified the gold-standard lists as the weakest assumption, noting that broad lists could make absolute performance a scoring artifact while preserving the internal control comparison. My concern extends this: the scoring mechanism itself, an 80% character-similarity fuzzy match, may be biased across arms because the treatment changes the linguistic register of responses. Control participants using internet search can copy exact disease names from search results, whereas LLM users more often paraphrase conversational advice. If so, the control-vs-LLM gap in condition identification could be inflated by measurement, not just by the breadth of the gold list. This is a concrete, testable threat to the strongest quantitative claim. I do not think it overturns the paper: the disposition null result, the internal randomization, and the robustness checks provide independent support for the broader conclusion that LLM assistance did not improve laypeople's decisions. However, the specific 'worse than control' condition result and the dramatic 94.9%-vs-34.5% knowledge-translation gap should be re-scored with blinded semantic judgment before the headline is taken at face value for policy. The manuscript is transparent about its methods and provides data and code, which makes this test feasible. My recommendation is therefore to retain the reader's CONDITIONAL verdict; the concern warrants a check but does not by itself justify rejection.","tokens_in":22232,"tokens_out":10377,"duration_ms":107034,"concrete_test":"Take all 2,400 free-text condition responses (or a stratified random sample of at least 300 per arm) and have two clinicians blinded to treatment assignment judge whether each response identifies any gold-standard condition, using semantic equivalence rather than character similarity. Compare the control-vs-LLM odds ratio under this semantic scoring with the odds ratio from the paper's fuzzy-match scoring. If the semantic scoring reproduces the control advantage and the LLM-user rate stays below about 35%, the concern is resolved; if the gap attenuates materially, the headline condition metric is partly a lexical scoring artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim that LLM-assisted participants were worse than control at identifying relevant conditions rests on the free-text condition metric (Methods, Scoring; Supplementary, Relevant condition ablations). A response is scored correct if any term fuzzy-matches at an 80% character-similarity threshold to any condition in the union of four physicians' differential lists. This is largely an orthographic match: it cannot recognize semantically correct lay paraphrases such as 'blood clot in the lung' for 'pulmonary embolism' or 'kidney infection' for 'pyelonephritis'. The control arm was instructed to use usual home resources, mostly internet search, where exact disease names are readily copyable; LLM users typically receive conversational paraphrases and must restate them in their own words. The two arms therefore differ systematically in the lexical form of their responses, so the observed control advantage (47.0% vs <34.5%) could be partly a scoring artifact rather than a true difference in comprehension. This is load-bearing because the headline 'clinical knowledge does not translate' is most strongly supported by the condition-identification gap, including the 94.9%-alone vs <34.5%-with-users contrast. The disposition result (no significant difference vs control) is not affected by this concern and still supports 'no benefit', but the stronger 'worse than control' claim for conditions requires the scoring assumption to hold.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a randomized controlled trial with 1,298 UK participants who were asked to assess one of ten physician-written medical scenarios. Participants were randomized to one of three LLM assistance arms (GPT-4o, Llama 3, Command R+) or a control arm that used usual home resources (mainly internet search). The study measured whether participants identified at least one condition from a physician-generated gold-standard list and whether they selected the correct disposition on a five-point acuity scale. The LLMs alone answered the same scenario questions correctly in 94.9% of cases for conditions and 56.3% for dispositions, but participants using the LLMs identified relevant conditions in at most 34.5% of responses (control: 47.0%) and showed no significant improvement in disposition accuracy relative to control. The paper further argues that performance on MedQA and on simulated-patient interactions does not predict these human-LLM interaction failures, and recommends human user testing before deployment.","tokens_in":22442,"tokens_out":6274,"duration_ms":67278,"significance":"If the findings are robust, the paper provides important evidence that high standalone performance of LLMs on medical knowledge tasks does not guarantee effective assistance for lay users. Strengths include a preregistration-style randomized design, stratified sampling, physician-generated scenarios and gold standards, a control group with realistic home resources, attrition analysis, public code and data, and multiple robustness checks. The finding that participants using LLMs were no better than control on disposition and worse on condition identification, despite the underlying LLMs performing well alone, is a practically consequential result for regulators and healthcare providers. However, the strength of the condition-identification claim depends on the validity of the fuzzy-matching scoring procedure, and the simulated-patient comparison rests on a very small number of simulation repetitions. These issues do not undermine the overall direction of the work but require attention before the headline claims can be accepted.","major_comments":[{"comment":"The central comparison for relevant conditions (Fig. 2b: 47.0% control vs at most 34.5% for each LLM arm, and Fig. 2: 94.9% LLM-alone vs below 34.5% with users) rests on an orthographic fuzzy-match rule. A response is scored as correct if it has at least 80% character similarity to any term on the physician-generated gold-standard list. This rule cannot recognize semantically equivalent lay paraphrases such as 'blood clot in the lung' for pulmonary embolism or 'kidney infection' for pyelonephritis. The control arm, instructed to use usual home resources, can copy exact disease names from internet search results, whereas LLM users receive conversational explanations and must restate them in their own words. The two arms are therefore not symmetric under this scoring rule. The threshold was tuned on 200 manually scored cases, but that sample was used to select the threshold and is not an independent validation of arm-specific lexical behavior; the reported 95.8% precision and recall do not rule out a systematic undercount of LLM-arm paraphrases. I ask the authors to re-score a random sample of responses from each arm with blinded clinician semantic judgment and to report whether the condition-identification difference between LLM users and control, and the LLM-alone versus with-user gap, persist under that scoring.","section":"Methods, Scoring; Supplementary 'Relevant condition ablations'"},{"comment":"The conclusion that simulated users do not predict human performance is based on only 10 simulation repetitions per model-scenario cell (300 simulated conversations in total). With 10 Bernoulli trials per cell, observing 0% or 100% accuracy in 26 of 30 scenarios is expected under substantial sampling variation, and the per-scenario regression coefficients have wide interval estimates (e.g., 0.20 ± 0.38 for Command R+ on disposition; -0.01 ± 0.51 for Command R+ on relevant conditions). The point estimates near zero are suggestive, but the evidence is too weak to support the strong claim that simulated interactions 'do not predict' human-LLM failures. A larger number of simulation repetitions, or reporting interval estimates for the correlations and a power analysis, would be needed to make this claim convincing.","section":"Methods, Simulated participants baseline; Fig. 4b"}],"minor_comments":[{"comment":"The analysis of conditions mentioned during user-LLM conversations uses GPT-4o to extract medical conditions from the transcripts, but this extraction is not validated. Since this extraction underlies the claim that LLMs suggested relevant conditions in 65.7-73.2% of conversations and that only 34.0% of suggestions were correct, a small validation study (e.g., comparing extraction against human annotation on a sample of transcripts) would strengthen the interaction-failure interpretation.","section":"Methods, User interactions; Fig. 3"},{"comment":"The MedQA subset contains only 236 questions, and the per-scenario samples are very small (e.g., tinnitus n=6, allergic rhinitis n=30). The comparison of benchmark accuracy to human experimental accuracy across scenarios would benefit from reporting uncertainty intervals for the per-scenario benchmark estimates, especially for scenarios with fewer than ten questions.","section":"Methods, Question-answering baseline"},{"comment":"The gold-standard condition lists are the union of four physicians' differentials and include some very broad items (e.g., 'cold' for allergic rhinitis; 'depression' for anaemia). Because participants are scored as correct if any fuzzy-matched term appears, the absolute condition-identification rates may be inflated by broad lists. I recommend reporting a sensitivity analysis that restricts scoring to red-flag conditions or to conditions that at least two physicians listed.","section":"Methods, Scenarios and Scoring; Supplementary Table 14"},{"comment":"The sentence 'The overall correct response rate of 43.0%±2.0% exceeds a random guessing baseline of 20%' is reported without a clear denominator; please clarify whether this is the average across all arms or the control arm, and state the corresponding test statistic.","section":"Results, Experimental performance"}],"recommendation":"major_revision","confidential_remarks":"The study is well executed and the data/code sharing is commendable. The main risk is the scoring artifact concern for the condition-identification outcome; if the authors can show with a blinded semantic re-scoring that the condition gap persists, the paper would be suitable for publication. The simulated-patient claim should also be tempered or supported with additional simulation runs. I do not see grounds for rejection, but the load-bearing condition-identification result needs validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the cleanest demonstration I've seen that a model's standalone medical accuracy doesn't carry into layperson use. 1,298 people, randomized, three models, ten physician-vetted scenarios; users were no better than control on disposition and worse on condition identification. Benchmarks and LLM-simulated users don't predict the interaction failure. That is a real, useful result for anyone building or regulating consumer health AI.\n\nThe paper earns its central claim. The RCT is well designed: stratified recruitment, scenario randomization, attrition check, public code and data. The interaction analysis, while not perfect, points to a concrete failure: LLMs offer reasonable conditions but users don't reliably carry them into final answers. I buy that expertise alone isn't sufficient.\n\nSoft spots are in the metrics, not the design. The condition-identification result rests on fuzzy string matching against physician-generated gold lists. The threshold was tuned on 200 manually scored cases with good precision and recall, which is better than nothing, but the tuning sample is small and the 80% character threshold cannot recognize a lay paraphrase like \"blood clot in the lung\" for pulmonary embolism. That matters because the control arm could copy exact disease names from search results, while LLM users get conversational paraphrases and have to restate them. So the \"worse than control\" condition gap may be partly a scoring artifact. The disposition result is immune to this and already shows no benefit, so the headline \"clinical knowledge does not translate\" survives. Still, the stronger condition claim should carry a caveat until someone re-scores a stratified sample with blinded semantic judgment. The same concern applies to the GPT-4o extraction of conditions from conversations, which is unvalidated and underpins the \"users only list 1.33 conditions\" analysis.\n\nThe simulated-user baseline is a reasonable comparison but only as good as the simulated patients; using GPT-4o for all of them is defensible, and the paper itself shows they don't track human variability, so I would treat that section as illustrative rather than a benchmark. The MedQA filtering is crude but not load-bearing.\n\nVerdict: this deserves a serious referee. I would ask for manual validation of the extraction step, sensitivity analyses on the gold-standard lists and fuzzy threshold, and a more guarded statement about condition identification. The main finding is important and largely holds.","headline":"Large RCT shows LLM assistance doesn't help laypeople; before taking the 'worse than control' condition result at face value, ask for a blinded semantic re-score.","tokens_in":23044,"tokens_out":2259,"would_cite":true,"duration_ms":23876,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that proficient LLMs provide no measurable benefit, and even some harm, when the public uses them for medical self-assessment.","keywords":["large language models","medical advice","human-computer interaction","randomized controlled trial","medical benchmarks","patient safety","AI-assisted decision making","health information"],"falsifier":"Take the same ten scenarios and re-score a fresh set of participants using a gold standard that includes only the single condition each scenario was designed to test, rather than the full union of differentials. If LLM-assisted participants then outperform the control group at naming that condition, the paper's 'no benefit' conclusion for condition identification would collapse; if they still do not, the interaction-failure explanation would be confirmed.","tokens_in":21998,"feed_emoji":"🏥","tokens_out":4138,"duration_ms":39583,"temperature":0.7,"pith_summary":"This paper reports a randomized controlled trial in which 1,298 members of the public used one of three large language models (GPT-4o, Llama 3, Command R+) to decide what to do about ten everyday medical scenarios. Tested alone, the models named a relevant condition in 94.9% of cases and chose the correct disposition in 56.3%. But the participants who used the models identified a relevant condition in at most 34.5% of cases, worse than the control group's 47.0%, and were no better than control at choosing the right disposition. The authors argue that the bottleneck is the human-model interaction, not the model's medical knowledge, and that neither medical licensing-style benchmarks nor simulated-patient interactions predict these failures. This matters because health services are considering deploying LLMs for direct public consultations.","feed_headline":"LLMs ace medical cases, but patients using them don't benefit","feed_subtitle":"In a 1,298-person trial, LLM users named fewer correct conditions than the unassisted control group.","key_machinery":"The study's central object is a randomized, between-subjects trial with four arms: three LLM-assisted arms and a control arm, all scored against physician-constructed gold standards. Ten scenarios were drafted by three doctors and a gold-standard list of relevant conditions was built from the union of differentials given by four additional doctors; the correct disposition was set by unanimous agreement. Two further comparison devices carry the argument: direct prompting of each LLM with the full scenario text (model-only performance), and a simulated-user condition where an LLM instance played the patient while another LLM provided assistance. These three layers isolate the effect of inserting a human user between a competent model and the task.","core_discovery":"The central discovery is that expert-level performance by an LLM in isolation does not translate into improved performance by a non-expert using that LLM, and can even reduce it. In the trial, LLM users were significantly less likely than the unassisted control group to name at least one condition from the physicians' gold-standard list (all three models, p<0.001), and their disposition accuracy was statistically indistinguishable from control. Analysis of the chat transcripts shows two points of breakdown: users often gave the model incomplete or misleading information, and even when the model suggested a correct condition, users frequently did not carry it into their final answer. These failures persist even though the same models score well above passing on a scenario-matched subset of MedQA, showing that knowledge benchmarks do not capture human-LLM interaction performance.","pith_inferences":["A concrete extension would be to vary the interface, for example by requiring structured questionnaires before the LLM responds, and test whether the condition-identification gap narrows; if it does, the failure is at information elicitation rather than comprehension or trust.","The condition-identification result may be sensitive to how the gold-standard list is scored; a strict scoring that only counts the single intended condition could change absolute numbers while preserving the control comparison, clarifying how much of the low rate is artifact.","The interaction-level finding suggests a latent failure mode for other high-stakes advisory domains, such as legal or financial advice: knowledge stored in the model is not the same as knowledge delivered through a chat exchange.","The weakness of simulated users as predictors points to a need for small but real human samples in safety evaluations, which regulators might mandate at far lower cost than is often imagined."],"forward_implications":["If this pattern holds, current medical licensing-style benchmarks overstate the readiness of LLMs for direct public use.","Simulated-patient evaluations that replace humans with LLMs can give misleadingly optimistic or falsely stable results.","Improving the base model's accuracy alone will not close the gap; effort must go into the interaction design, including information elicitation and conveying recommendations.","LLM users may underestimate the acuity of serious conditions, which carries asymmetric risk compared with overestimation.","Public-facing medical LLMs will need to be proactive in requesting missing information, rather than passively responding to whatever the user says."],"supporting_citations":[{"why":"Provides the MedQA benchmark used for the question-answering baseline comparison.","marker":"[4]"},{"why":"Shows that AI assistance did not improve radiologists' X-ray reading, motivating the design of human-model assistance trials.","marker":"[8]"},{"why":"Demonstrates that physicians assisted by LLMs barely outperformed unassisted physicians and both trailed the LLM alone, a pattern the paper extends to the public.","marker":"[9]"},{"why":"Supplies the 5-shot prompting method and medical knowledge benchmark framing for evaluating LLMs.","marker":"[21]"},{"why":"Provides the simulated-patient prompting approach that the paper adapts for its simulated-user baseline.","marker":"[25]"},{"why":"A randomized trial showing LLMs did not improve physicians' diagnostic reasoning, which the paper cites as prior evidence that expert knowledge does not automatically transfer to users.","marker":"[27]"}],"fun_headline_variants":["LLMs ace med exams but flop with real patients","Medical AI fails when humans ask for help","Expert LLMs, zero patient benefit in trial","AI knows diagnoses but can't get them across","Ask an LLM for medical help? Study says no benefit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The condition-identification results rest on the gold-standard lists of relevant conditions being the right target; these lists are the union of four physicians' free-text differentials, so if they are too broad, the low absolute scores could partly reflect a scoring artifact, although the comparison with the control group remains internally valid.","fun_headline_variants_meta":{"raw":{"variants":["LLMs ace med exams but flop with real patients","Medical AI fails when humans ask for help","Expert LLMs, zero patient benefit in trial","AI knows diagnoses but can't get them across","Ask an LLM for medical help? Study says no benefit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2730,"prompt_tokens":928,"completion_tokens":1802,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1727}},"tokens_in":544,"tokens_out":1802,"duration_ms":13641,"temperature":1.0,"reasoning_tokens":1727,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:06:01.012507+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same ten scenarios and re-score a fresh set of participants using a gold standard that includes only the single condition each scenario was designed to test, rather than the full union of differentials. If LLM-assisted participants then outperform the control group at naming that condition, the paper's 'no benefit' conclusion for condition identification would collapse; if they still do not, the interaction-failure explanation would be confirmed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MedQA benchmark used for the question-answering baseline comparison."},{"cited_title":"& Salz, T","cited_arxiv_id":null,"evidence_quote":"Shows that AI assistance did not improve radiologists' X-ray reading, motivating the design of human-model assistance trials."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 5-shot prompting method and medical knowledge benchmark framing for evaluating LLMs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the simulated-patient prompting approach that the paper adapts for its simulated-user baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A randomized trial showing LLMs did not improve physicians' diagnostic reasoning, which the paper cites as prior evidence that expert knowledge does not automatically transfer to users."}],"review_version":1}