{"id":"ac29b09f-ac3e-4b5d-90ac-a0c44117d0ae","arxiv_id":"2505.08902","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A commentary argues that medical research should stop running evanescent LLM-versus-human comparisons and instead study human-LLM collaboration, supported by a literature review showing rapid model turnover and tiny expert cohorts.","lead":"This communication argues that comparing each new LLM to a small group of human experts is too slow and unstable to guide patient care. It calls on medical researchers to study how clinicians and LLMs can work together.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central performance-gain claim is asserted, not demonstrated: its own cited RCT (Goh et al., ref 8) found LLM-alone outperformed physicians with LLM access, and refs 35-40 provide no controlled human+LLM comparison.","rationale":"The reader's weakest assumption matches the main vulnerability. I agree, but would sharpen it: the problem is not merely absence of an experiment in this paper; the paper's own cited evidence (ref 8) provides a controlled counterexample, and the references offered for the collaboration premise (refs 35-40) are not controlled comparisons. This makes the strong form of the claim internally inconsistent with cited material. However, the paper is a position/communication piece; a conditional verdict that requires rewording 'demonstrate' to 'argue', adding explicit task-dependence, and citing direct human+LLM evidence where it exists would be appropriate. The concern therefore reinforces rather than overturns the reader's conditional verdict.","tokens_in":7236,"tokens_out":4899,"duration_ms":50939,"concrete_test":"Perform a two-part check. (1) For each of refs 35-40, classify the study design: controlled comparison with arms (human-alone, LLM-alone, human+LLM), technical method without a human arm, or opinion/review. If no cited study includes a human+LLM arm, then the 'largest gains' claim is unsupported by the cited evidence. (2) Extract the primary outcome from Goh et al. (ref 8, JAMA Netw Open 2024): compare diagnostic reasoning scores across arms. If the LLM-alone arm is statistically best, then the paper must either state the claim as task-dependent or provide positive evidence for tasks where human augmentation reverses the ordering.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is stated in Section 'Bridging the Gap for Enhanced Performance': 'The augmentation of LLM output with human input... is how the largest performance gains from LLMs can be realized,' supported by ref 38. If this premise fails, the title's performance-gain claim and the paper's main redirect both collapse. The paper provides no measurement of any human+LLM system against human-alone or LLM-alone baselines. The cited support does not fill the gap: refs 35-36 address context limitations and contrastive decoding, ref 37 is a fine-tuning review, ref 38 is a roadmap/opinion, and ref 39 documents increased collaboration, not performance gains. More seriously, the paper's own Introduction cites ref 8 (Goh et al., JAMA Netw Open 2024), a randomized trial in which LLM alone outperformed both clinicians with LLM access and clinicians without on diagnostic reasoning. That is direct evidence that, at least for some decision tasks, adding human input did not produce the largest gain. The abstract says 'we demonstrate,' but what the paper actually establishes is a reasonable critique of comparison-study methodology; the positive claim that collaboration yields the largest gains is an unstated hypothesis, not a demonstrated result. Because the central recommendation depends on this unsupported and partially contradicted premise, the strong form of the claim overreaches.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This communication argues that medical research on large language models (LLMs) should move away from \"LLMs versus human experts\" comparison studies and toward studying human-LLM collaboration. The paper presents a compiled set of 1,567 studies comparing LLMs to human experts, notes the rapid obsolescence of specific model versions, the variability of the term \"expert,\" the small sizes of expert comparator groups, and the geographic concentration of authors in English-speaking countries. It then calls for rigorous, device-style clinical evaluation of LLMs and a global LLM registry, and concludes that \"the largest performance gains from LLMs\" will come from augmenting LLM output with human input rather than replacing humans.","tokens_in":7472,"tokens_out":5094,"duration_ms":46928,"significance":"The paper addresses a genuine and timely problem: much of the current evaluation literature is indeed hard to accumulate because model iterations are discontinued and comparator groups are often tiny and ill-defined. The recommendation to complement, rather than simply replace, human expertise with LLMs is a reasonable research priority. However, the empirical contribution is a literature count with underreported methodology, and the affirmative claim of \"largest performance gains\" is not demonstrated; the strongest direct evidence cited (ref 8) runs against the claim in at least one task. The manuscript's value is as a position statement, not as a demonstration, and it needs reframing and a reproducible review protocol to support its conclusions.","major_comments":[{"comment":"The central claim that \"The augmentation of LLM output with human input instead of competing against and replacing human input is how the largest performance gains from LLMs can be realized\" is load-bearing for the title and conclusion, but it is supported only by citations to roadmap/opinion work (refs 35-40), none of which provides a controlled comparison of human+LLM against LLM-alone or human-alone. The paper's own Introduction (p. 2) cites Goh et al. (ref 8), a randomized clinical trial in which LLM-alone outperformed physicians with LLM access on diagnostic reasoning; that evidence directly undercuts the universal form of the claim. The manuscript must either supply such a comparison or explicitly re-label the statement as an untested hypothesis and temper the abstract and title accordingly.","section":"Bridging the Gap for Enhanced Performance (p. 6)"},{"comment":"The literature review is not reproducible. The paper states that search terms were developed with a medical librarian and reports 1,567 identified studies, but provides no search strings, no list of databases beyond the specific PubMed counts for model names, no date ranges, and no inclusion/exclusion criteria. The statement that \"we randomly selected 100 studies and found a total of 59 studies which compared LLMs to human experts after filtering\" omits the filtering procedure, and the raw data are not deposited. This prevents verification of the empirical claims (e.g., the 54% discontinuation figure in Figure 1, the exponential growth in Figure 2) and weakens the paper's stated basis for \"we demonstrate.\"","section":"Introduction (pp. 2-3)"},{"comment":"The abstract declares \"we demonstrate\" that there is a need for human-LLM collaboration, and the title asserts \"Performance Gains of LLMs With Humans.\" What the manuscript actually establishes is that the comparison literature exists, is growing, and has methodological limitations (model turnover, small expert panels, geographic skew). The positive claim that collaboration yields the largest performance gains is asserted rather than shown; no experiment, meta-analysis, or derived performance comparison is presented. The paper should be reframed as a perspective or hypothesis piece, with the title and abstract revised to match, unless an actual comparative evaluation is added.","section":"Abstract and Conclusion (pp. 1, 6-7)"}],"minor_comments":[{"comment":"The legend \"Proportional of studies\" should read \"Proportion of studies.\"","section":"Figure 3 (p. 4)"},{"comment":"The phrase \"we demonstrate\" overstates the nature of a commentary supported by a non-reproducible literature count; \"we argue\" or \"we propose\" would be more accurate.","section":"Abstract (p. 1)"},{"comment":"The label \"Exponential increase\" implies a fitted curve, but no model fit or uncertainty is reported; if an exponential trend is intended, provide the fitted equation and confidence interval, otherwise change the label to \"Increase over time.\"","section":"Figure 2 (p. 3)"},{"comment":"Items such as \"conclude that 'we should be cautious'\" are rhetorical and shift the tone of an otherwise scholarly argument; reconsider whether this style is appropriate for the journal.","section":"Introduction bullet list (p. 5)"},{"comment":"References 5 and 7 appear to be the published and preprint versions of the same clinical text summarization study; citing both is confusing, and the preprint could be removed.","section":"References (pp. 8-9)"},{"comment":"The sentence \"LLMs are not sentient, which is dangerous for practice in healthcare when it comes to administering painkillers or anesthetists\" is vague; the causal pathway from lack of sentience to danger in pain management should be clarified or supported by a specific clinical incident.","section":"LLMs Lack Emotional Understanding (p. 5)"}],"recommendation":"major_revision","confidential_remarks":"The paper is closer to a commentary than an empirical study; if the journal publishes perspective pieces, the framing issues can be fixed. The authors' affiliation with MIT Critical Data and the use of the hive-learning framework (ref 40) as corroboration for the collaboration agenda is a possible perceived-conflict point; the editor may wish to have the authors make that connection explicit or broaden the supporting literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a short position paper, not a study, but it has one piece of genuinely useful new synthesis: the compiled numbers on how fast LLM comparison studies decay. Fifty-four percent of models compared to human experts in their sample are already discontinued, nearly 80% of studies compare against fewer than ten experts, and the authorship is heavily skewed to the US and China. Those numbers are a real contribution to the critique of this literature, and the paper does well to point out that 'expert' is used inconsistently and that publication lag makes snapshot comparisons obsolete.\n\nThe soft spot is the title and abstract. They say 'we demonstrate' performance gains from human-LLM collaboration, but the paper contains no experiment, no data on human+LLM systems, and no comparison against baselines. The performance-gain claim is asserted via reference 38, which is a roadmap/opinion piece. Worse, the paper's own introduction cites Goh et al. (ref 8), a randomized trial in which the LLM alone outperformed clinicians both with and without LLM access on diagnostic reasoning. That doesn't disprove the collaboration hypothesis for other tasks, but it does mean the strongest form of the claim is unsupported and partially contradicted by the paper's own evidence. The honest version of the paper is: 'LLM-versus-expert studies are methodologically weak and quickly outdated; we should move toward longitudinal, population-specific human-in-the-loop evaluation.' That is a reasonable agenda, but the title promises something stronger.\n\nAlso, the methodology of the literature counts is thin. Search terms, dates, inclusion criteria, and raw data are not given, so the 54% and 80% figures can't be independently checked. That's fixable with a supplementary file but matters because the numbers are the paper's only novel contribution.\n\nThe citation pattern is fine; refs 35-40 do not support the empirical claim, but they're used in a way consistent with their actual content. There's self-citation of the MIT Critical Data 'hive learning' framework that is incidental rather than load-bearing.\n\nWho's this for? People working on clinical LLM evaluation who need a compact argument for why head-to-head benchmark comparisons are a poor use of effort. It could be useful in a reading group. It is not a research contribution in the performance-gain sense. I'd send it to review only if the authors revise the title, add methods detail and data release, and soften the central claim to what they actually show. As is, it's a borderline case that a serious referee could improve.","headline":"A useful empirical sketch of how transient LLM-versus-human studies are, wrapped in a title that promises performance evidence the paper doesn't deliver.","tokens_in":8045,"tokens_out":2123,"would_cite":false,"duration_ms":20132,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Medical AI research should replace 'LLM versus human expert' studies with 'LLM with human expert' studies.","keywords":["LLM clinical deployment","human-AI collaboration","medical AI evaluation","model-iteration-agnostic research","patient safety","human expert comparisons","AI regulation in medicine"],"falsifier":"A prospective randomized trial comparing, on the same clinical task, (i) a current LLM alone, (ii) clinicians alone, and (iii) clinician-LLM pairs, repeated after each major model release: the central claim collapses if the paired arm does not beat both solo arms, or if the gain disappears when the model is updated.","tokens_in":7040,"feed_emoji":"🤝","tokens_out":8476,"duration_ms":79671,"temperature":0.7,"pith_summary":"This paper argues that the current boom of 'LLM versus human expert' studies in medicine is the wrong research paradigm: the models being tested are often discontinued before studies publish, the comparison groups are small and ill-defined, and the framing ignores that LLMs will be used with clinicians, not instead of them. The authors identify 1,567 such studies and show that most compare against fewer than ten people called 'experts,' with authors concentrated in English-speaking countries. They call for a shift to 'LLMs with human experts' research, with LLMs evaluated like medical devices through randomized Level 1 trials, uncertainty-aware endpoints, and a global open database of model performance across populations. The reason to care is that the choice of research paradigm shapes regulation, deployment, and patient safety.","feed_headline":"Pair doctors with LLMs instead of pitting them against each other","feed_subtitle":"Every model tested against 'experts' goes stale fast; the paper argues human-AI teams are the path to real gains.","key_machinery":"The central move is a reframing from 'what' questions to 'how' questions: instead of asking what a specific, temporary LLM can do better than a small group of humans, ask how the model should be inserted into the existing healthcare workflow. The proposed mechanism is a human final layer in the LLM pipeline, where clinicians fill the gaps the model cannot cover — context, empathy, local knowledge, and attention weighting for specific tasks — while LLMs provide scale and speed. The paper also specifies supporting machinery: randomized Level 1 clinical trial designs with endpoints for demographic parity and uncertainty communication, a global open database tracking model performance across populations, reporting guidelines, and conclusions that are deliberately agnostic to which model generation produced them.","core_discovery":"The central claim is that research built on comparing LLMs to a loosely defined group of human experts cannot guide clinical deployment, because the technologies turn over faster than the literature: in the authors' review, 54 percent of the models compared with human experts had already been discontinued or rebranded, and one widely used model had 443 literature records before being retired after about a year. Most comparisons are also statistically thin, with nearly 80 percent of sampled studies comparing LLMs to fewer than ten human experts, and the 'expert' group ranges from physicians to average test takers. The paper's positive thesis is that the largest performance gains will come from collaboration, not replacement: humans should supply emotional understanding, local knowledge, and contextual judgment while LLMs supply speed and scale, so the research agenda should move from 'what can this temporary model do' to 'how can this model be integrated, monitored, and regulated for safe patient care.'","pith_inferences":["The same obsolescence argument applies outside medicine: in law, finance, and engineering, model versions also churn faster than peer review, so solo 'model versus expert' benchmarks may be replaced by workflow-integration studies there too.","A testable extension of the paper's thesis is that human-plus-LLM team performance will be more stable across model generations than solo LLM performance, because the human layer absorbs a model's context and empathy failures.","The paper's own numbers imply that any study naming a specific model version has a finite shelf life; an editor could require authors to state how long their conclusions are expected to hold.","If the collaborative agenda is correct, the practical bottleneck will shift from model capability to interface design, workflow design, and clinician training."],"forward_implications":["Comparative 'LLM versus human' evaluations should be used only when no alternative answers a specific clinical question, and then run under rigorous human-evaluation guidelines.","LLM evaluation in medicine should be model-iteration-agnostic, so that findings state principles that survive the next model release instead of describing a discontinued snapshot.","LLMs should be held to the evidence standard of medical devices, with randomized Level 1 trials covering demographic parity, uncertainty communication, and deployment design.","Research priorities should include a global open database of LLM performance across populations, safeguards against data poisoning, reporting guidelines, and accounting for carbon cost and access inequities.","The field should concentrate on 'how' questions — integration, monitoring, regulation, and human-LLM interaction — rather than benchmark 'what' questions."],"supporting_citations":[{"why":"A randomized clinical trial showing LLMs alone outperformed humans with and without LLM access on diagnostic reasoning; it motivates the replacement question the paper argues against.","marker":"[8]"},{"why":"Evidence that LLMs struggle to integrate contextual features, establishing the gap that human input is meant to fill.","marker":"[35]"},{"why":"An example of model-side architectural enhancement for contextual understanding, used to contrast with human-side supplementation.","marker":"[36]"},{"why":"Widely cited source for fine-tuning foundational LLMs, supporting the idea that human input can weight attention for specific tasks.","marker":"[37]"},{"why":"The paper's key citation for the claim that augmenting LLM output with human input, rather than replacing humans, realizes the largest performance gains.","marker":"[38]"},{"why":"Provides the human-evaluation framework the paper recommends when comparisons are necessary.","marker":"[18]"},{"why":"Supplies the LLM reporting guideline listed among the needed next steps for the field.","marker":"[21]"},{"why":"Shows peer-reviewed publication takes months to years, underpinning the argument that comparison studies are too slow for rapidly updated models.","marker":"[11]"},{"why":"Supports the idea that authors and reviewers without model-specific knowledge create a 'blind leading the blind' dynamic in the literature.","marker":"[12]"}],"fun_headline_variants":["LLM + human teams: the real clinical win","Don't compare LLMs to doctors—combine them","Why LLM vs. expert is a race to obsolescence","Human-AI symbiosis beats model-versus-human tests","Collaboration, not competition, is the LLM edge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that adding a human to an LLM workflow yields the largest performance gains assumes the human layer fixes the model's context, empathy, and local-knowledge weaknesses without introducing new, serious error modes.","fun_headline_variants_meta":{"raw":{"variants":["LLM + human teams: the real clinical win","Don't compare LLMs to doctors—combine them","Why LLM vs. expert is a race to obsolescence","Human-AI symbiosis beats model-versus-human tests","Collaboration, not competition, is the LLM edge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000682,"raw_usage":{"total_tokens":3072,"prompt_tokens":898,"completion_tokens":2174,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2092}},"tokens_in":514,"tokens_out":2174,"duration_ms":16196,"temperature":1.0,"reasoning_tokens":2092,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:44:34.841039+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective randomized trial comparing, on the same clinical task, (i) a current LLM alone, (ii) clinicians alone, and (iii) clinician-LLM pairs, repeated after each major model release: the central claim collapses if the paired arm does not beat both solo arms, or if the gain disappears when the model is updated.","supporting_citations":[{"cited_title":"Knowledge-Empowered, Collaborative, and Co-Evolving AI Models: The Post-LLM Roadmap","cited_arxiv_id":null,"evidence_quote":"The paper's key citation for the claim that augmenting LLM output with human input, rather than replacing humans, realizes the largest performance gains."},{"cited_title":"Pandemic publishing: Medical journals strongly speed up their publication process for COVID-19","cited_arxiv_id":null,"evidence_quote":"Shows peer-reviewed publication takes months to years, underpinning the argument that comparison studies are too slow for rapidly updated models."}],"review_version":1}