REVIEW 3 major objections 6 minor 40 references
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Medical AI research should replace 'LLM versus human expert' studies with 'LLM with human expert' studies.
desk verdict A useful empirical sketch of how transient LLM-versus-human studies are, wrapped in a title that promises performance evidence the paper doesn't deliver. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central move is a reframing from 'what' questions to 'how' questions: instead of asking what a specific, temporary LLM can do better than a small group of humans, ask how the model should be inserted into the existing healthcare workflow. The proposed mechanism is a human final layer in the LLM pipeline, where clinicians fill the gaps the model cannot cover — context, empathy, local knowledge, and attention weighting for specific tasks — while LLMs provide scale and speed. The paper also specifies supporting machinery: randomized Level 1 clinical trial designs with endpoints for demographic parity and uncertainty communication, a global open database tracking model performance across populations, reporting guidelines, and conclusions that are deliberately agnostic to which model generation produced them.
What would settle it
A prospective randomized trial comparing, on the same clinical task, (i) a current LLM alone, (ii) clinicians alone, and (iii) clinician-LLM pairs, repeated after each major model release: the central claim collapses if the paired arm does not beat both solo arms, or if the gain disappears when the model is updated.
Extended reading notes
Core claim
The central claim is that research built on comparing LLMs to a loosely defined group of human experts cannot guide clinical deployment, because the technologies turn over faster than the literature: in the authors' review, 54 percent of the models compared with human experts had already been discontinued or rebranded, and one widely used model had 443 literature records before being retired after about a year. Most comparisons are also statistically thin, with nearly 80 percent of sampled studies comparing LLMs to fewer than ten human experts, and the 'expert' group ranges from physicians to average test takers. The paper's positive thesis is that the largest performance gains will come from collaboration, not replacement: humans should supply emotional understanding, local knowledge, and contextual judgment while LLMs supply speed and scale, so the research agenda should move from 'what can this temporary model do' to 'how can this model be integrated, monitored, and regulated for safe patient care.'
Load-bearing premise
The claim that adding a human to an LLM workflow yields the largest performance gains assumes the human layer fixes the model's context, empathy, and local-knowledge weaknesses without introducing new, serious error modes.
Editorial extensions
If this is right
- Comparative 'LLM versus human' evaluations should be used only when no alternative answers a specific clinical question, and then run under rigorous human-evaluation guidelines.
- LLM evaluation in medicine should be model-iteration-agnostic, so that findings state principles that survive the next model release instead of describing a discontinued snapshot.
- LLMs should be held to the evidence standard of medical devices, with randomized Level 1 trials covering demographic parity, uncertainty communication, and deployment design.
- Research priorities should include a global open database of LLM performance across populations, safeguards against data poisoning, reporting guidelines, and accounting for carbon cost and access inequities.
- The field should concentrate on 'how' questions — integration, monitoring, regulation, and human-LLM interaction — rather than benchmark 'what' questions.
Reading between the lines
- The same obsolescence argument applies outside medicine: in law, finance, and engineering, model versions also churn faster than peer review, so solo 'model versus expert' benchmarks may be replaced by workflow-integration studies there too.
- A testable extension of the paper's thesis is that human-plus-LLM team performance will be more stable across model generations than solo LLM performance, because the human layer absorbs a model's context and empathy failures.
- The paper's own numbers imply that any study naming a specific model version has a finite shelf life; an editor could require authors to state how long their conclusions are expected to hold.
- If the collaborative agenda is correct, the practical bottleneck will shift from model capability to interface design, workflow design, and clinician training.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This communication argues that medical research on large language models (LLMs) should move away from "LLMs versus human experts" comparison studies and toward studying human-LLM collaboration. The paper presents a compiled set of 1,567 studies comparing LLMs to human experts, notes the rapid obsolescence of specific model versions, the variability of the term "expert," the small sizes of expert comparator groups, and the geographic concentration of authors in English-speaking countries. It then calls for rigorous, device-style clinical evaluation of LLMs and a global LLM registry, and concludes that "the largest performance gains from LLMs" will come from augmenting LLM output with human input rather than replacing humans.
Significance. The paper addresses a genuine and timely problem: much of the current evaluation literature is indeed hard to accumulate because model iterations are discontinued and comparator groups are often tiny and ill-defined. The recommendation to complement, rather than simply replace, human expertise with LLMs is a reasonable research priority. However, the empirical contribution is a literature count with underreported methodology, and the affirmative claim of "largest performance gains" is not demonstrated; the strongest direct evidence cited (ref 8) runs against the claim in at least one task. The manuscript's value is as a position statement, not as a demonstration, and it needs reframing and a reproducible review protocol to support its conclusions.
major comments (3)
- [Bridging the Gap for Enhanced Performance (p. 6)] The central claim that "The augmentation of LLM output with human input instead of competing against and replacing human input is how the largest performance gains from LLMs can be realized" is load-bearing for the title and conclusion, but it is supported only by citations to roadmap/opinion work (refs 35-40), none of which provides a controlled comparison of human+LLM against LLM-alone or human-alone. The paper's own Introduction (p. 2) cites Goh et al. (ref 8), a randomized clinical trial in which LLM-alone outperformed physicians with LLM access on diagnostic reasoning; that evidence directly undercuts the universal form of the claim. The manuscript must either supply such a comparison or explicitly re-label the statement as an untested hypothesis and temper the abstract and title accordingly.
- [Introduction (pp. 2-3)] The literature review is not reproducible. The paper states that search terms were developed with a medical librarian and reports 1,567 identified studies, but provides no search strings, no list of databases beyond the specific PubMed counts for model names, no date ranges, and no inclusion/exclusion criteria. The statement that "we randomly selected 100 studies and found a total of 59 studies which compared LLMs to human experts after filtering" omits the filtering procedure, and the raw data are not deposited. This prevents verification of the empirical claims (e.g., the 54% discontinuation figure in Figure 1, the exponential growth in Figure 2) and weakens the paper's stated basis for "we demonstrate."
- [Abstract and Conclusion (pp. 1, 6-7)] The abstract declares "we demonstrate" that there is a need for human-LLM collaboration, and the title asserts "Performance Gains of LLMs With Humans." What the manuscript actually establishes is that the comparison literature exists, is growing, and has methodological limitations (model turnover, small expert panels, geographic skew). The positive claim that collaboration yields the largest performance gains is asserted rather than shown; no experiment, meta-analysis, or derived performance comparison is presented. The paper should be reframed as a perspective or hypothesis piece, with the title and abstract revised to match, unless an actual comparative evaluation is added.
minor comments (6)
- [Figure 3 (p. 4)] The legend "Proportional of studies" should read "Proportion of studies."
- [Abstract (p. 1)] The phrase "we demonstrate" overstates the nature of a commentary supported by a non-reproducible literature count; "we argue" or "we propose" would be more accurate.
- [Figure 2 (p. 3)] The label "Exponential increase" implies a fitted curve, but no model fit or uncertainty is reported; if an exponential trend is intended, provide the fitted equation and confidence interval, otherwise change the label to "Increase over time."
- [Introduction bullet list (p. 5)] Items such as "conclude that 'we should be cautious'" are rhetorical and shift the tone of an otherwise scholarly argument; reconsider whether this style is appropriate for the journal.
- [References (pp. 8-9)] References 5 and 7 appear to be the published and preprint versions of the same clinical text summarization study; citing both is confusing, and the preprint could be removed.
- [LLMs Lack Emotional Understanding (p. 5)] The sentence "LLMs are not sentient, which is dangerous for practice in healthcare when it comes to administering painkillers or anesthetists" is vague; the causal pathway from lack of sentience to danger in pain management should be clarified or supported by a specific clinical incident.
Circularity Check
No circularity: no derivation chain or fitted inputs; self-citation is rhetorical and non-load-bearing.
full rationale
This manuscript is an opinion/position communication rather than a quantitative derivation. It contains no equations, no fitted parameters, no model trained on a subset and validated on a related superset, and no prediction that is equivalent by construction to its input. The central claim—that human-LLM collaboration rather than human-LLM comparison will yield the largest performance gains—is asserted in the section 'Bridging the Gap for Enhanced Performance' and supported by external citations (especially ref 38), not derived from any data produced in the paper. The cited reference 40 ('Village mentoring and hive learning: The MIT Critical Data experience') may involve overlapping authors and is used rhetorically at the end of that section ('the future of village mentoring and hive learning'), but it is not load-bearing: the core recommendation does not reduce to that citation, and the paper would stand or fall on the external evidence for collaborative LLM use regardless of ref 40. The fact that the paper says 'we demonstrate' without actually providing a controlled human+LLM comparison is a weakness in evidence and support, and the cited RCT in the Introduction (ref 8) may even complicate the collaboration premise, but that is a correctness/overreach concern, not a circularity loop. No step in the paper's argument is defined in terms of its conclusion, renames a known result, or imports a uniqueness theorem from the authors' prior work. Under the hard rules requiring a quoted reduction or a fitted-parameter-as-prediction before flagging circularity, the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption LLMs are not sentient and lack general empathy unless guided, which makes autonomous LLM use unsafe in healthcare.
- domain assumption Human-LLM collaboration realizes the largest performance gains.
- domain assumption Discontinuation of an LLM model makes comparison studies obsolete.
- domain assumption LLMs should be regulated like medical devices.
Cite this review
Pith. "Pith review of Performance Gains of LLMs With Humans in a World of LLMs Versus Humans." pith.science (2026). https://pith.science/paper/346P3RWQ
@misc{pith2026250508902,
author = {Pith},
title = {Pith review of: Performance Gains of LLMs With Humans in a World of LLMs Versus Humans},
year = {2026},
howpublished = {\url{https://pith.science/paper/346P3RWQ}},
note = {Machine review of arXiv:2505.08902}
}
read the original abstract
Currently, a considerable research effort is devoted to comparing LLMs to a group of human experts, where the term "expert" is often ill-defined or variable, at best, in a state of constantly updating LLM releases. Without proper safeguards in place, LLMs will threaten to cause harm to the established structure of safe delivery of patient care which has been carefully developed throughout history to keep the safety of the patient at the forefront. A key driver of LLM innovation is founded on community research efforts which, if continuing to operate under "humans versus LLMs" principles, will expedite this trend. Therefore, research efforts moving forward must focus on effectively characterizing the safe use of LLMs in clinical settings that persist across the rapid development of novel LLM models. In this communication, we demonstrate that rather than comparing LLMs to humans, there is a need to develop strategies enabling efficient work of humans with LLMs in an almost symbiotic manner.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Emergent Abilities of Large Language Models
Wei J, Tay Y, Bommasani R, et al. Emergent Abilities of Large Language Models. Published online October 26, 2022. doi:10.48550/arXiv.2206.07682
-
[2]
Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. Dagan A, ed. PLOS Digit Health. 2023;2(2):e0000198. doi:10.1371/journal.pdig.0000198
-
[3]
Bhayana R, Krishna S, Bleakney RR. Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations. Radiology. 2023;307(5):e230582. doi:10.1148/radiol.230582
-
[4]
Chatbot vs Medical Student Performance on Free-Response Clinical Reasoning Examinations
Strong E, DiGiammarino A, Weng Y, et al. Chatbot vs Medical Student Performance on Free-Response Clinical Reasoning Examinations. JAMA Intern Med. 2023;183(9):1028. doi:10.1001/jamainternmed.2023.2909
-
[5]
Adapted large language models can outperform medical experts in clinical text summarization
Van Veen D, Van Uden C, Blankemeier L, et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat Med. 2024;30(4):1134-1142. doi:10.1038/s41591-024-02855-5
-
[6]
Cascella M, Semeraro F, Montomoli J, Bellini V, Piazza O, Bignami E. The Breakthrough of Large Language Models Release for Medical Applications: 1-Year Timeline and Perspectives. J Med Syst. 2024;48(1):22. doi:10.1007/s10916-024-02045-3
-
[7]
Clinical Text Summarization: Adapting Large Language Models Can Outperform Human Experts
Veen DV, Uden CV, Blankemeier L, et al. Clinical Text Summarization: Adapting Large Language Models Can Outperform Human Experts. Published online October 30, 2023. doi:10.21203/rs.3.rs-3483777/v1
-
[8]
Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial
Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open. 2024;7(10):e2440969. doi:10.1001/jamanetworkopen.2024.40969
arXiv 2024
Show all 40 references
-
[9]
OpenAI o1-Preview vs
Temsah MH, Jamal A, Alhasan K, Temsah AA, Malki KH. OpenAI o1-Preview vs. ChatGPT in Healthcare: A New Frontier in Medical AI Reasoning. Cureus. Published online October 1, 2024. doi:10.7759/cureus.70640
2024 doi
-
[10]
Robustness of evidence reported in preprints during peer review
Nelson L, Ye H, Schwenn A, Lee S, Arabi S, Hutchins BI. Robustness of evidence reported in preprints during peer review. The Lancet Global Health. 2022;10(11):e1684-e1687. doi:10.1016/S2214- 109X(22)00368-0
2022 doi
-
[11]
Pandemic publishing: Medical journals strongly speed up their publication process for COVID-19
Horbach SPJM. Pandemic publishing: Medical journals strongly speed up their publication process for COVID-19. Quantitative Science Studies. 2020;1(3):1056-1067. doi:10.1162/qss_a_00076
2020 doi
-
[12]
The future landscape of large language models in medicine
Clusmann J, Kolbinger FR, Muti HS, et al. The future landscape of large language models in medicine. Commun Med. 2023;3(1):1-8. doi:10.1038/s43856-023-00370-1 8
2023 doi
- [13]
- [14]
- [15]
- [16]
-
[17]
Medical large language models are vulnerable to data-poisoning attacks
Alber DA, Yang Z, Alyakin A, et al. Medical large language models are vulnerable to data-poisoning attacks. Nat Med. Published online January 8, 2025. doi:10.1038/s41591-024-03445-1
2025 doi
-
[18]
A framework for human evaluation of large language models in healthcare derived from literature review
Tam TYC, Sivarajkumar S, Kapoor S, et al. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digit Med. 2024;7(1):258. doi:10.1038/s41746-024- 01258-7
2024 doi
- [19]
- [20]
-
[21]
The TRIPOD-LLM reporting guideline for studies using large language models
Gallifant J, Afshar M, Ameen S, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med. 2025;31(1):60-69. doi:10.1038/s41591-024-03425-5
2025 doi
-
[22]
Reconciling the contrasting narratives on the environmental impact of large language models
Ren S, Tomlinson B, Black RW, Torrance AW. Reconciling the contrasting narratives on the environmental impact of large language models. Sci Rep. 2024;14(1):26310. doi:10.1038/s41598-024- 76682-6
2024 doi
-
[23]
Emotional intelligence and patient-centred care
Birks YF, Watt IS. Emotional intelligence and patient-centred care. J R Soc Med. 2007;100(8):368-374
2007
-
[24]
Physiciansʼ Empathy and Clinical Outcomes for Diabetic Patients: Academic Medicine
Hojat M, Louis DZ, Markham FW, Wender R, Rabinowitz C, Gonnella JS. Physiciansʼ Empathy and Clinical Outcomes for Diabetic Patients: Academic Medicine. 2011;86(3):359-364. doi:10.1097/ACM.0b013e3182086fe1
2011 doi
-
[25]
Large Language Models and Empathy: Systematic Review
Sorin V, Brin D, Barash Y, et al. Large Language Models and Empathy: Systematic Review. J Med Internet Res. 2024;26:e52597. doi:10.2196/52597
2024 doi
-
[26]
Emotional intelligence of Large Language Models
Wang X, Li X, Yin Z, Wu Y, Liu J. Emotional intelligence of Large Language Models. Journal of Pacific Rim Psychology. 2023;17:18344909231213958. doi:10.1177/18344909231213958
2023 doi
-
[27]
Third-party evaluators perceive AI as more compassionate than expert humans
Ovsyannikova D, De Mello VO, Inzlicht M. Third-party evaluators perceive AI as more compassionate than expert humans. Commun Psychol. 2025;3(1):4. doi:10.1038/s44271-024-00182-6
2025 doi
- [28]
- [29]
-
[30]
Ethical and Practical Issues with Opioids in Life-Limiting Illness
Fine RL. Ethical and Practical Issues with Opioids in Life-Limiting Illness. Baylor University Medical Center Proceedings. 2007;20(1):5-12. doi:10.1080/08998280.2007.11928223
2007
-
[31]
Which Humans? Published online September 22, 2023
Atari M, Xue MJ, Park PS, Blasi DE, Henrich J. Which Humans? Published online September 22, 2023. 9 doi:10.31234/osf.io/5b26t
2023 doi
- [32]
- [33]
-
[34]
Generative language models exhibit social identity biases
Hu T, Kyrychenko Y, Rathje S, Collier N, Van Der Linden S, Roozenbeek J. Generative language models exhibit social identity biases. Nat Comput Sci. 2024;5(1):65-75. doi:10.1038/s43588-024-00741-1
2024 doi
- [35]
- [36]
-
[37]
The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities
Parthasarathy VB, Zafar A, Khan A, Shahid A. The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities. Published online October 30, 2024. doi:10.48550/arXi...
-
[38]
Knowledge-Empowered, Collaborative, and Co-Evolving AI Models: The Post-LLM Roadmap
Wu F, Shen T, Bäck T, et al. Knowledge-Empowered, Collaborative, and Co-Evolving AI Models: The Post-LLM Roadmap. Engineering. 2025;44:87-100. doi:10.1016/j.eng.2024.12.008
2025 doi
-
[39]
Academic collaboration on large language model studies increases overall but varies across disciplines
Li L, Dinh L, Hu S, Hemphill L. Academic collaboration on large language model studies increases overall but varies across disciplines. Published online 2024. doi:10.48550/ARXIV.2408.04163
-
[40]
Village mentoring and hive learning: The MIT Critical Data experience
Cosgriff CV, Charpignon M, Moukheiber D, et al. Village mentoring and hive learning: The MIT Critical Data experience. iScience. 2021;24(6):102656. doi:10.1016/j.isci.2021.102656
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.