Pith. sign in

REVIEW 2 major objections 5 minor 272 references

Generative AI in Medicine

T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This review argues that generative AI's medical usefulness is gated less by what models can generate than by unresolved problems of consent, privacy, transparency, hallucination, usability, equity, real-world evaluation, and accountability.

desk verdict Useful stakeholder-organized synthesis; add a search-strategy paragraph and temper two headline stats before calling it comprehensive. read the letter →

arxiv 2412.10337 v2 pith:RXAPXYRW submitted 2024-12-13 cs.LG cs.AIcs.CYcs.HC

classification cs.LGcs.AIcs.CYcs.HC
keywords generativeAImedicinehealthcarelargelanguagemodelsdiffusionclinicaltrialshealthequityreal-worldevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper surveys the landscape of generative AI in medicine and argues that the technology's real-world benefit is currently gated by a small set of deployment problems rather than by model capability. It organizes use cases by the people who would use them—clinicians, patients, clinical trial organizers, researchers, and trainees—and shows that each group already has working pilots: note drafting, query answering, trial screening, hypothesis generation, and case creation for education. The same capability that makes these pilots possible also creates new risks, including memorized patient data, plausible misinformation, biased outputs, and failures to follow clinical guidelines. The authors conclude that realizing the potential requires progress on eight challenge areas—informed consent, privacy and security, transparency and interpretability, hallucination mitigation, interface usability, equity, real-world evaluation, and accountability—and they point to open research directions in each.

What carries the argument

The paper's organizing device is a stakeholder-based taxonomy that sorts use cases into five columns—clinicians, patients, clinical trial organizers, researchers, and trainees—and a challenge list of eight rows that cut across all of them. This matrix does the argumentative work: it converts an unwieldy collection of pilots and benchmarks into a structured map showing which use cases have early quantitative evidence, which challenges are shared across stakeholders, and where the open research questions cluster. The taxonomy is paired with a technical primer on the three model families (text, image, and text-plus-image) that expresses, in plain terms, how each model type fits the medical tasks it is proposed for.

What would settle it

A systematic literature search with explicit inclusion criteria that surfaces a major use case or challenge absent from this review would weaken the comprehensiveness claim; more directly, if the headline pilot results the paper repeats—the 88 percent differential-diagnosis rate, the 90 percent reduction in eligibility checks, and the 42 percent time saving—fail to reproduce when tested on larger, more diverse patient populations, the paper's 'promising' emphasis on those use cases would not hold.

Watch

Extended reading notes

Core claim

The paper's central contention is that generative AI in medicine is genuinely promising but not yet dependable: current models can already include the correct diagnosis in a differential in roughly the same proportion as clinicians, cut the number of eligibility criteria a clinician must check for trial enrollment by about 90 percent, and draft patient-facing messages and notes that clinicians find usable. Yet the same models fail to follow diagnostic guidelines up to 36 percent of the time, leak private training data when prompted adversarially, and reproduce or exaggerate demographic stereotypes. These paired observations drive the paper's organizing claim: the gap between what medical generative AI can do and what should be deployed is a gap in consent, privacy, transparency, usability, equity, evaluation, and accountability, not a gap in raw generation ability. The paper frames this as a roadmap for research and policy rather than a prediction about any single model.

Load-bearing premise

The review's claim to be comprehensive rests on the unstated assumption that the cited studies are a representative sample of the broader literature on generative AI in medicine, since the paper describes no systematic search strategy or inclusion criteria.

Editorial extensions

If this is right

  • Health systems and regulators should concentrate near-term effort on deployment infrastructure—consent workflows, privacy-preserving training, transparency reporting, and evaluation protocols—rather than waiting for larger models.
  • Each of the five stakeholder groups will need its own validation standard; benchmark accuracy alone will not suffice, since the paper shows that guideline adherence, information ordering, and patient comprehension can diverge from correctness.
  • Synthetic data generation will become a standard tool for privacy-preserving dataset construction and fairness improvement, but only if paired with auditing for inherited biases.
  • Clinician workflows will shift from drafting to supervising: the paper's evidence suggests models can produce drafts and differentials, but significant editing and verification remain necessary.
  • Medical education will adopt generated clinical vignettes and feedback at scale, with diversity auditing as a required guardrail to avoid amplifying demographic over-indexing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the challenge list is correct, the field's near-term progress will be measured less by new model releases than by the maturity of evaluation frameworks that test guideline adherence, information-order sensitivity, and patient comprehension—the paper implies this but does not say it as a prediction.
  • The stakeholder taxonomy is extensible: adding groups such as medical coders, public-health officials, or informal caregivers would likely surface additional use cases and challenges the review does not cover.
  • A testable extension: models tuned on medical corpora will show smaller advantages over general models on real patient questions than on benchmark exams, because the paper's evaluation discussion suggests benchmark sets misrepresent real query distributions.
  • The paper's emphasis on accountability and over-reliance implies that liability rules and interface design, not technical accuracy, will likely determine whether these tools actually improve patient outcomes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is a narrative review of generative AI in medicine. It organizes applications by five stakeholder groups—clinicians, patients, clinical trial organizers, researchers, and trainees—and then surveys eight challenge areas (informed consent, privacy/security, transparency/interpretability, hallucinations, usability, equity, real-world evaluation, accountability). The paper's central claim, stated in the abstract and Section 1, is that these use cases and challenges constitute a comprehensive overview of the field, and that the challenges must be addressed before the potential can be realized.

Significance. If the claims hold, the paper's value is primarily organizational and pedagogical: it provides a concise, well-referenced map of a rapidly growing literature and links use cases to concrete open problems. The stakeholder-based taxonomy is a reasonable and useful frame for researchers entering the area, and the challenge sections, particularly those on usability, equity, and real-world evaluation, synthesize recent literature more effectively than most competing reviews. The paper is also careful to attribute quantitative claims to specific studies, and it explicitly names its own funding and a co-author's company affiliation, which is appropriate for a review. The main risk is that the 'comprehensive' label is stronger than the unstated literature-selection process can guarantee.

major comments (2)
  1. [Abstract and Section 1] The abstract and Section 1 claim that the paper provides a 'comprehensive overview' of generative AI in medicine, but no search strategy, inclusion criteria, or screening process is described anywhere in the manuscript. Because the review uses specific quantitative results to motivate each use case (e.g., 88% differential-diagnosis accuracy in Section 2.1; 90% eligibility-check reduction in Section 2.3), the selection of cited studies is load-bearing: if the chosen papers overrepresent pilot studies or particular research networks, the comprehensiveness claim and the relative emphasis among use cases would not be reproducible by readers. Please add a short methods paragraph (databases, date range, inclusion/exclusion criteria, and how the stakeholder taxonomy was populated) or, if this is intended as a non-systematic narrative review, temper the 'comprehensive' claim in the abstract and Section 1 and add a scope-limitation statement.
  2. [Section 2.1 and Section 2.3] The two headline quantitative claims—88% differential-diagnosis inclusion versus clinicians' 96% (ref 61) and a 90% reduction in manually checked eligibility criteria (ref 107)—are stated without the study context necessary to interpret them. In both cases, the underlying studies are single, non-randomized evaluations on curated or small-scale data, and the numbers likely depend on specific model versions, prompts, and outcome definitions; the manuscript should report these conditions or hedge the findings as preliminary pilot results rather than general benchmarks. The 42% time reduction (ref 109) also lacks any sample-size or study-design information. Without this context, the review risks overstating the maturity of these use cases and misdirecting readers' expectations.
minor comments (5)
  1. [Section 3.7] The sentence 'common medical LLM benchmarks use questions from exams (256, 257), clinical guidebooks (184, 187), or research papers (258)' cites references 184 and 187, which are studies of cancer-treatment information rather than clinical guidebooks; please verify the intended references or rephrase the category.
  2. [Section 2.4] In the description of reference 127, the phrase 'outperforms supervised baselines' should be 'outperform supervised baselines' to agree with the plural subject 'hypotheses.'
  3. [Section 3.4] The sentence beginning 'Policymakers can, for example, encourage greater transparency...' is long and would benefit from splitting to improve readability.
  4. [Throughout] The manuscript uses 'generative AI' and 'generative interfaces' interchangeably; consider defining 'generative interface' on first use and using the terms consistently.
  5. [Figures 1 and 2] The figures are referenced but not included in the provided text; please ensure the final figures match the bullet lists in Sections 2 and 3 and are legible in print.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; this is a narrative review with no derivation chain, no fitted parameters, and no self-cited uniqueness argument.

full rationale

The paper is a narrative review that organizes existing work by stakeholder group and lists challenges; it contains no equations, no fitted parameters, and no first-principles derivation whose output could be equivalent to its input by construction. The headline statistics cited to motivate use cases (e.g., 88% differential-diagnosis accuracy, 90% reduction in manually checked eligibility criteria) are drawn from external papers (refs 61 and 107), not from the authors' own prior work, and the review does not fit or modify those numbers. The authors cite several of their own papers (refs 26, 58, 121, 225, 226, 229, 231, 233, 235, 239, 266), but these citations support background points or specific sub-claims, and the central taxonomy and challenge list are independently supported by a broad external literature. There is no self-citation that is load-bearing, no imported uniqueness theorem, and no ansatz smuggled in via citation. The absence of a described search strategy is a legitimate methodological limitation, but it concerns representativeness and correctness, not circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a review, so the ledger contains no fitted parameters and no invented entities. The load-bearing premises are domain assumptions: the cited primary studies are accurately reported and representative, the non-systematic literature selection adequately supports the comprehensive scope claim, and the stakeholder taxonomy is a complete partition of the space. These assumptions enter at the Introduction and throughout Section 2, where specific summary statistics from single studies are elevated to representative findings.

assumptions (3)
  • domain assumption The cited primary studies are accurately reported and their results are representative of the broader literature on generative AI in medicine.
    The review's summary statistics, such as 88 percent differential coverage, 91 percent screening accuracy, and 90 percent eligibility-criteria reduction, are taken from single cited studies without independent verification or meta-analytic framing. Introduced throughout Sections 2.1 to 2.5.
  • domain assumption The non-systematic selection of papers adequately supports the 'comprehensive' scope claim.
    No search strategy, inclusion criteria, or registration is provided in Section 1 or Section 2, yet the abstract claims a comprehensive overview; the scope claim rests entirely on this unstated selection process.
  • domain assumption The stakeholder taxonomy (clinicians, patients, trial organizers, researchers, trainees) is a complete and useful partition of medical generative AI use cases.
    The framing in Section 2 excludes some stakeholder groups such as hospital administrators, payers, and regulators, and Section 1.1 acknowledges excluding data modalities like physiological signals and molecular graphs; completeness is assumed rather than demonstrated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generative AI in Medicine." pith.science (2026). https://pith.science/paper/RXAPXYRW

@misc{pith2026241210337,
  author       = {Pith},
  title        = {Pith review of: Generative AI in Medicine},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RXAPXYRW}},
  note         = {Machine review of arXiv:2412.10337}
}
read the original abstract

The increased capabilities of generative AI have dramatically expanded its possible use cases in medicine. We provide a comprehensive overview of generative AI use cases for clinicians, patients, clinical trial organizers, researchers, and trainees. We then discuss the many challenges -- including maintaining privacy and security, improving transparency and interpretability, upholding equity, and rigorously evaluating models -- which must be overcome to realize this potential, and the open research directions they give rise to.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

272 extracted references · 54 canonical work pages

  1. [1]

    Tai-Seale M, Baxter SL, Vaida F, Walker A, Sitapati AM, et al. 2024. AI-Generated Draft Replies Integrated Into Health Records and Physicians’ Electronic Communication. JAMA Network Open 7(4):e246565–e246565

  2. [2]

    Zakka C, Cho J, Fahed G, Shad R, Moor M, et al. 2024. Almanac copilot: Towards autonomous electronic health record navigation. arXiv preprint arXiv:2405.07896

  3. [3]

    Garcia P, Ma SP, Shah S, Smith M, Jeong Y, et al. 2024. Artificial intelligence–generated draft replies to patient inbox messages. JAMA Network Open 7(3):e243201–e243201

  4. [4]

    Kambhamettu H, Metaxa D, Johnson K, Head A. 2024. Explainable Notes: Examining How to Unlock Meaning in Medical Notes with Interactivity and Artificial Intelligence . In Proceedings of the CHI Conference on Human Factors in Computing Systems , CHI ’24. New York, NY, USA: Association for Computing Machinery www.annualreviews.org • Generative AI in Medicine 17

  5. [5]

    Mannhardt N, Bondi-Kelly E, Lam B, O’Connell C, Asiedu M, et al. 2024. Impact of large language model assistance on patients reading clinical notes: A mixed-methods study. arXiv preprint arXiv:2401.09637

  6. [6]

    Tierney AA, Gayre G, Hoberman B, Mattern B, Ballesca M, et al. 2024. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catalyst Innova- tions in Care Delivery 5(3):CAT–23

  7. [7]

    Marshall IJ, Nye B, Kuiper J, Noel-Storr A, Marshall R, et al. 2020. Trialstreamer: A living, automatically updated database of clinical trial reports. Journal of the American Medical Informatics Association 27(12):1903–1912

  8. [8]

    Feldman J, Hochman KA, Guzman BV, Goodman A, Weisstuch J, Testa P. 2024. Scaling Note Quality Assessment Across an Academic Medical Center with AI and GPT-4. NEJM Catalyst 5(5):CAT.23.0283

Show all 272 references
  1. [9]

    Liu S, Wright AP, Mccoy AB, Huang SS, Genkins JZ, et al. 2024. Using large language model to guide patients to create efficient and comprehensive clinical care message. Journal of the American Medical Informatics Association :ocae142

  2. [10]

    Wang Z, Xiao C, Sun J. 2023. AutoTrial: Prompting Language Models for Clinical Trial Design. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, ed. H Bouamor, J Pino, K Bali, pp. 12461–12472, pp. 12461–12472. Singapore: Association for C...

  3. [11]

    White RD, Peng T, Sripitak P, Johansen AR, Snyder M. 2023. CliniDigest: A Case Study in Large Language Model Based Large-Scale Summarization of Clinical Trial Descriptions . In Proceedings of the 2023 ACM Conference on Information Technology for Social Good , pp. 396–402

  4. [12]

    Guo E, Gupta M, Deng J, Park YJ, Paget M, Naugler C. 2024. Automated Paper Screening for Clinical Reviews Using Large Language Models: Data Analysis Study. Journal of Medical Internet Research 26:e48996

  5. [13]

    Ktena I, Wiles O, Albuquerque I, Rebuffi SA, Tanno R, et al. 2024. Generative models improve fairness of medical classifiers under distribution shifts. Nature Medicine 30(4):1166–1173

  6. [14]

    Bakkum MJ, Hartjes MG, Pi¨ et JD, Donker EM, Likic R, et al. 2024. Using artificial intelligence to create diverse and inclusive medical case vignettes for education. British Journal of Clinical Pharmacology 90(3):640–648

  7. [15]

    Dupont PE, Nelson BJ, Goldfarb M, Hannaford B, Menciassi A, et al. 2021. A decade retro- spective of medical robotics research from 2010 to 2020. Science robotics 6(60):eabi8017

  8. [16]

    Kononenko I. 2001. Machine learning for medical diagnosis: history, state of the art and perspective. Artificial Intelligence in medicine 23(1):89–109

  9. [17]

    Razzak MI, Imran M, Xu G. 2020. Big data analytics for preventive medicine. Neural Com- puting and Applications 32(9):4417–4451

  10. [18]

    Kruse CS, Smith B, Vanderlinden H, Nealand A. 2017. Security techniques for the electronic health records. Journal of medical systems 41:1–9

  11. [19]

    Fern´ andez-Alem´ an JL, Se˜ nor IC, Lozoya P´AO, Toval A. 2013. Security and privacy in elec- tronic health records: A systematic literature review. Journal of biomedical informatics 46(3):541–562

  12. [20]

    Campanella P, Lovato E, Marone C, Fallacara L, Mancuso A, et al. 2016. The impact of electronic health records on healthcare quality: a systematic review and meta-analysis. The European Journal of Public Health 26(1):60–64

  13. [21]

    Graber ML, Siegal D, Riah H, Johnston D, Kenyon K. 2019. Electronic health record–related events in medical malpractice claims. Journal of patient safety 15(2):77–85

  14. [22]

    Ancker JS, Edwards A, Nosal S, Hauser D, Mauer E, et al. 2017. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC medical informatics and decision making 17:1–9

  15. [23]

    Rosenberg N. 1982. Learning by using. Inside the black box: Technology and economics :120– 18 Shanmugam et al. 140

  16. [24]

    Bishop CM, Nasrabadi NM. 2006. Pattern recognition and machine learning , vol. 4. Springer

  17. [25]

    Kaplan J, McCandlish S, Henighan T, Brown TB, Chess B, et al. 2020. Scaling Laws for Neural Language Models. arXiv:2001.08361 [cs, stat] ArXiv: 2001.08361

  18. [26]

    Movva R, Balachandar S, Peng K, Agostini G, Garg N, Pierson E. 2024. Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Ling...

  19. [27]

    Abbaspourazad S, Elachqar O, Miller AC, Emrani S, Nallasamy U, Shapiro I. 2023. Large-scale Training of Foundation Models for Wearable Biosignals

  20. [28]

    Beaini D, Huang S, Cunha JA, Li Z, Moisescu-Pareja G, et al. 2023. Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task Datasets

  21. [29]

    McKeen K, Oliva L, Masood S, Toma A, Rubin B, Wang B. 2024. ECG-FM: An Open Elec- trocardiogram Foundation Model

  22. [30]

    Bond-Taylor S, Leach A, Long Y, Willcocks CG. 2021. Deep Generative Modelling: A Compar- ative Review of V AEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models

  23. [31]

    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, et al. 2017. Attention is all you need. Advances in neural information processing systems 30

  24. [32]

    Jurafsky D. 2000. Speech and language processing

  25. [33]

    Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, et al. 2022. Training language models to follow instructions with human feedback. ArXiv:2203.02155 [cs]

  26. [34]

    Bolton E, Venigalla A, Yasunaga M, Hall D, Xiong B, et al. 2024. BioMedLM: A 2.7B Param- eter Language Model Trained On Biomedical Text. ArXiv:2403.18421 [cs]

  27. [35]

    Chen Z, Cano AH, Romanou A, Bonnet A, Matoba K, et al. 2023. MEDITRON-70B: Scaling Medical Pretraining for Large Language Models. ArXiv:2311.16079 [cs]

  28. [37]

    Jin D, Pan E, Oufattole N, Weng WH, Fang H, Szolovits P. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Ex- ams. Applied Sciences 11(14):6421Number: 14 Publisher: Multidisciplinary Digital Publishing Institute

  29. [38]

    Fleming SL, Lozano A, Haberkorn WJ, Jindal JA, Reis EP, et al. 2023. MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records. ArXiv:2308.14089 [cs]

  30. [39]

    Rasmy L, Xiang Y, Xie Z, Tao C, Zhi D. 2021. Med-BERT: pretrained contextualized em- beddings on large-scale structured electronic health records for disease prediction. npj Digital Medicine 4(1):1–13Publisher: Nature Publishing Group

  31. [40]

    Hill BL, Emami M, Nori VS, Cordova-Palomera A, Tillman RE, Halperin E. 2023. CHIRon: A Generative Foundation Model for Structured Sequential Medical Data

  32. [41]

    Ferruz N, Schmidt S, H¨ ocker B. 2022. Protgpt2 is a deep unsupervised language model for protein design. Nature communications 13(1):4348

  33. [42]

    Rives A, Meier J, Sercu T, Goyal S, Lin Z, et al. 2021. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences 118(15):e2016239118Publisher: Proceedings of the National Acade...

  34. [43]

    Nguyen E, Poli M, Durrant MG, Thomas A W, Kang B, et al. 2024. Sequence modeling and design from molecular to genome scale with Evo. Pages: 2024.02.27.582234 Section: New Results

  35. [44]

    Yang L, Zhang Z, Song Y, Hong S, Xu R, et al. 2023. Diffusion Models: A Comprehensive Survey of Methods and Applications. ACM Comput. Surv. 56(4):105:1–105:39 www.annualreviews.org • Generative AI in Medicine 19

  36. [45]

    Dhariwal P, Nichol A. 2021. Diffusion Models Beat GANs on Image Synthesis. ArXiv:2105.05233 [cs, stat]

  37. [46]

    Kazerouni A, Aghdam EK, Heidari M, Azad R, Fayyaz M, et al. 2023. Diffusion models in medical imaging: A comprehensive survey. Medical Image Analysis 88:102846

  38. [47]

    Sun S, Goldgof G, Butte A, Alaa AM. 2023. Aligning Synthetic Medical Images with Clini- cal Knowledge using Human Feedback. Advances in Neural Information Processing Systems 36:13408–13428

  39. [48]

    Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, et al. 2021. Learning Transferable Vi- sual Models From Natural Language Supervision . In Proceedings of the 38th International Conference on Machine Learning , pp. 8748–8763. PMLR. ISSN: 2640-3498

  40. [49]

    Chambon PJM, Bluethgen C, Langlotz C, Chaudhari A. 2022. Adapting Pretrained Vision- Language Foundational Models to Medical Imaging Domains

  41. [50]

    Li J, Li D, Savarese S, Hoi S. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Machine Learning , pp. 19730–19742. PMLR. ISSN: 2640-3498

  42. [51]

    Thawakar OC, Shaker AM, Mullappilly SS, Cholakkal H, Anwer RM, et al. 2024. XrayGPT: Chest Radiographs Summarization using Large Medical Vision-Language Models . In Pro- ceedings of the 23rd Workshop on Biomedical Natural Language Processing , ed. D Demner- Fushman, S Ananiado...

  43. [52]

    Bazi Y, Rahhal MMA, Bashmal L, Zuair M. 2023. Vision–Language Model for Visual Ques- tion Answering in Medical Imagery. Bioengineering 10(3):380Number: 3 Publisher: Multidis- ciplinary Digital Publishing Institute

  44. [53]

    Li C, Wong C, Zhang S, Usuyama N, Liu H, et al. 2023. LLaV A-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. Advances in Neural Information Processing Systems 36:28541–28564

  45. [54]

    Aiken LH, Lasater KB, Sloane DM, Pogue CA, Fitzpatrick Rosenbaum KE, et al. 2023. Physi- cian and Nurse Well-Being and Preferred Interventions to Address Burnout in Hospital Prac- tice: Factors Associated With Turnover, Outcomes, and Patient Safety. JAMA Health Forum 4(7):e231809

  46. [55]

    Saag HS, Shah K, Jones SA, Testa PA, Horwitz LI. 2019. Pajama time: working after work in the electronic health record. Journal of general internal medicine 34:1695–1696

  47. [56]

    Heer J. 2019. Agency plus automation: Designing artificial intelligence into interactive systems. Proceedings of the National Academy of Sciences 116(6):1844–1850

  48. [57]

    Mamykina L, Vawdrey DK, Stetson PD, Zheng K, Hripcsak G. 2012. Clinical documenta- tion: composition or synthesis? Journal of the American Medical Informatics Association 19(6):1025–1031

  49. [58]

    Jiang S, Shen S, Agrawal M, Lam B, Kurtzman N, et al. 2023. Conceptualizing machine learn- ing for dynamic information retrieval of electronic health record notes . In Machine Learning for Healthcare Conference, pp. 343–359. PMLR

  50. [59]

    Young CC, Enichen E, Rivera C, Auger CA, Grant N, et al. 2024. Diagnostic Accuracy of a Custom Large Language Model on Rare Pediatric Disease Case Reports. American Journal of Medical Genetics Part A n/a(n/a):e63878

  51. [60]

    Kanjee Z, Crowe B, Rodman A. 2023. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA

  52. [61]

    Levine DM, Tuwani R, Kompa B, Varma A, Finlayson SG, et al. 2023. The Diagnostic and Triage Accuracy of the GPT-3 Artificial Intelligence Model

  53. [62]

    R ´ ıos-Hoyo A, Shan NL, Li A, Pearson AT, Pusztai L, Howard FM. 2024. Evaluation of large language models as a diagnostic aid for complex medical cases. Frontiers in Medicine 11

  54. [63]

    Olmo Jd, Logro˜ no J, Masc ´ ıas C, Mart ´ ınez M, Isla J. 2024. Assessing DxGPT: Diagnosing Rare Diseases with Various Large Language Models 20 Shanmugam et al

  55. [64]

    Kanjee Z, Crowe B, Rodman A. 2023. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330(1):78–80

  56. [65]

    Zhou J, He X, Sun L, Xu J, Chen X, et al. 2024. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nature Communications 15(1):5649

  57. [66]

    Thawkar O, Shaker A, Mullappilly SS, Cholakkal H, Anwer RM, et al. 2023. XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

  58. [67]

    Moor M, Huang Q, Wu S, Yasunaga M, Zakka C, et al. 2023. Med-Flamingo: a Multimodal Medical Few-shot Learner

  59. [68]

    Lin Z, Zhang D, Tao Q, Shi D, Haffari G, et al. 2023. Medical visual question answering: A survey. Artificial Intelligence in Medicine 143:102611

  60. [69]

    Reese JT, Danis D, Caufield JH, Groza T, Casiraghi E, et al. 2024. On the limitations of large language models in clinical diagnosis. medRxiv :2023.07.13.23292613

  61. [70]

    Hager P, Jungmann F, Holland R, Bhagat K, Hubrecht I, et al. 2024. Evaluation and mitiga- tion of the limitations of large language models in clinical decision-making. Nature Medicine 30(9):2613–2622

  62. [71]

    Ahmed A, Chandra S, Herasevich V, Gajic O, Pickering BW. 2011. The effect of two different electronic health record user interfaces on intensive care provider task load, errors of cognition, and performance. Critical care medicine 39(7):1626–1634

  63. [72]

    Murray L, Gopinath D, Agrawal M, Horng S, Sontag D, Karger DR. 2021. Medknowts: unified documentation and information retrieval for electronic health records . In The 34th Annual ACM Symposium on User Interface Software and Technology , pp. 1169–1183

  64. [73]

    Zheng K, Padman R, Johnson MP, Diamond HS. 2009. An interface-driven analysis of user interactions with an electronic health records system. Journal of the American Medical Infor- matics Association 16(2):228–237

  65. [74]

    Sackett DL, Rosenberg WMC. 1995. On the need for evidence-based medicine. Journal of Public Health 17(3):330–334

  66. [75]

    Bastian H, Glasziou P, Chalmers I. 2010. Seventy-five trials and eleven systematic reviews a day: how will we ever keep up? PLoS medicine 7(9):e1000326

  67. [76]

    Zakka C, Chaurasia A, Shad R, Dalal AR, Kim JL, et al. 2023. Almanac: Retrieval-Augmented Language Models for Clinical Medicine. ArXiv:2303.01229 [cs]

  68. [77]

    Lee K, Paek H, Huang LC, Hilton CB, Datta S, et al. 2024. SEETrials: Leveraging Large Language Models for Safety and Efficacy Extraction in Oncology Clinical Trials

  69. [78]

    Chewning B, Bylund CL, Shah B, Arora NK, Gueguen JA, Makoul G. 2012. Patient preferences for shared decisions: a systematic review. Patient education and counseling 86(1):9–18

  70. [79]

    Adler RF, Morales P, Sotelo J, Magasi S. 2022. Developing an mhealth app for empowering cancer survivors with disabilities: Co-design study. JMIR Formative Research 6(7):e37706

  71. [80]

    Noack EM, Schulze J, M¨ uller F. 2021. Designing an app to overcome language barriers in the delivery of emergency medical services: participatory development process. JMIR mHealth and uHealth 9(4):e21586

  72. [81]

    Danieli M, Ciulli T, Mousavi SM, Riccardi G. 2021. A conversational artificial intelligence agent for a mental health care app: Evaluation study of its participatory design. JMIR Formative Research 5(12):e30053

  73. [82]

    Martin-Hammond A, Vemireddy S, Rao K, et al. 2019. Exploring older adults’ beliefs about the use of intelligent assistants for consumer health information management: A participatory design study. JMIR aging 2(2):e15381

  74. [83]

    Amante DJ, Hogan TP, Pagoto SL, English TM, Lapane KL. 2015. Access to care and use of the internet to search for health information: results from the us national health interview survey. Journal of medical Internet research 17(4):e106

  75. [84]

    Thapa DK, Visentin DC, Kornhaber R, West S, Cleary M. 2021. The influence of online health information on health decisions: A systematic review. Patient education and counseling 104(4):770–784 www.annualreviews.org • Generative AI in Medicine 21

  76. [85]

    Vanessa Choy, Sara Martin, Ashley Lumpkin. 2024. Can we rely on generative AI for healthcare information? | Ipsos

  77. [86]

    Alex Montero, Grace Sparks, Marley Presiado, Liz Hamel. 2024. KFF Health Misinformation Tracking Poll: Health and Election Issues on TikTok | KFF

  78. [87]

    Hersh W. 2024. Search still matters: information retrieval in the era of generative ai. Journal of the American Medical Informatics Association :ocae014

  79. [88]

    Grossman L V, Feiner SK, Mitchell EG, Creber RMM. 2018. Leveraging patient-reported out- comes using data visualization. Applied clinical informatics 9(03):565–575

  80. [89]

    Zhao JY, Song B, Anand E, Schwartz D, Panesar M, et al. 2017. Barriers, facilitators, and solutions to optimal patient portal and personal health record use: a systematic review of the literature. In AMIA annual symposium proceedings , vol. 2017, pp. 1913. American Medical Inf...

  81. [90]

    Grossman L V, Choi SW, Collins S, Dykes PC, O’Leary KJ, et al. 2018. Implementation of acute care patient portals: recommendations on utility and use from six early adopters. Journal of the American Medical Informatics Association 25(4):370–379

  82. [91]

    Grossman L V, Masterson Creber RM, Benda NC, Wright D, Vawdrey DK, Ancker JS. 2019. Interventions to increase patient portal use in vulnerable populations: a systematic review. Journal of the American Medical Informatics Association 26(8-9):855–870

  83. [92]

    Warren LR, Harrison M, Arora S, Darzi A. 2019. Working with patients and the public to design an electronic health record interface: a qualitative mixed-methods study. BMC medical informatics and decision making 19:1–8

  84. [93]

    Kambhamettu H, Metaxa D, Johnson K, Head A. 2024. Explainable Notes: Examining How to Unlock Meaning in Medical Notes with Interactivity and Artificial Intelligence . In Proceedings of the CHI Conference on Human Factors in Computing Systems , pp. 1–19

  85. [94]

    Luo L, Vairavamurthy J, Zhang X, Kumar A, Ter-Oganesyan RR, et al. 2024. ReXplain: Translating Radiology into Patient-Friendly Video Reports

  86. [95]

    Basu C, Vasu R, Yasunaga M, Yang Q. 2023. Med-EASi: finely annotated dataset and models for controllable simplification of medical texts . In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of ...

  87. [96]

    Mirza FN, Tang OY, Connolly ID, Abdulrazeq HA, Lim RK, et al. 2024. Using chatgpt to facilitate truly informed medical consent. NEJM AI 1(2):AIcs2300145

  88. [97]

    Shadbolt C, Naufal E, Bunzli S, Price V, Rele S, et al. 2023. Analysis of Rates of Completion, Delays, and Participant Recruitment in Randomized Clinical Trials in Surgery.JAMA Network Open 6(1):e2250996

  89. [98]

    Ross JS, Mocanu M, Lampropulos JF, Tse T, Krumholz HM. 2013. Time to Publication Among Completed Clinical Trials. JAMA Internal Medicine 173(9):825–828

  90. [99]

    Zarin DA, Tse T, Williams RJ, Rajakannan T. 2017. Update on Trial Registration 11 Years after the ICMJE Policy Was Established. New England Journal of Medicine 376(4):383–391

  91. [100]

    Getz KA, Stergiopoulos S, Short M, Surgeon L, Krauss R, et al. 2016. The Impact of Protocol Amendments on Clinical Trial Performance and Cost. Therapeutic Innovation & Regulatory Science 50(4):436–441

  92. [101]

    Fogel DB. 2018. Factors associated with clinical trials that fail and opportunities for improving the likelihood of success: A review. Contemporary Clinical Trials Communications 11:156–164

  93. [102]

    Ghim JL, Ahn S. 2023. Transforming clinical trials: the emerging roles of large language models. Translational and Clinical Pharmacology 31(3):131–138

  94. [103]

    Park J, Fang Y, Ta C, Zhang G, Idnay B, et al. 2024. Criteria2Query 3.0: Leveraging generative large language models for clinical trial eligibility query generation. Journal of Biomedical Informatics 154:104649

  95. [104]

    Lai H, Ge L, Sun M, Pan B, Huang J, et al. 2024. Assessing the Risk of Bias in Randomized 22 Shanmugam et al. Clinical Trials With Large Language Models. JAMA Network Open 7(5):e2412687

  96. [105]

    Datta S, Lee K, Paek H, Manion FJ, Ofoegbu N, et al. 2024. AutoCriteria: a generalizable clinical trial eligibility criteria extraction system powered by large language models. Journal of the American Medical Informatics Association 31(2):375–385

  97. [106]

    Yuan C, Ryan PB, Ta C, Guo Y, Li Z, et al. 2019. Criteria2Query: a natural language inter- face to clinical databases for cohort definition. Journal of the American Medical Informatics Association 26(4):294–305

  98. [107]

    Hamer DMd, Schoor P, Polak TB, Kapitan D. 2023. Improving Patient Pre-screening for Clinical Trials: Assisting Physicians with Large Language Models. ArXiv:2304.07396 [cs]

  99. [108]

    Wornow M, Lozano A, Dash D, Jindal J, Mahaffey KW, Shah NH. 2024. Zero-Shot Clinical Trial Patient Matching with LLMs

  100. [109]

    Jin Q, Wang Z, Floudas CS, Chen F, Gong C, et al. 2024. Matching Patients to Clinical Trials with Large Language Models. ArXiv :arXiv:2307.15051v4

  101. [110]

    Beattie J, Neufeld S, Yang D, Chukwuma C, Gul A, et al. 2024. Utilizing Large Language Models for Enhanced Clinical Trial Matching: A Study on Automation in Patient Screening. Cureus 16(5):e60044

  102. [111]

    McCann S, Campbell M, Entwistle V. 2013. Recruitment to clinical trials: a meta-ethnographic synthesis of studies of reasons for participation. Journal of Health Services Research & Policy 18(4):233–241

  103. [112]

    Skea ZC, Newlands R, Gillies K. 2019. Exploring non-retention in clinical trials: a meta- ethnographic synthesis of studies reporting participant reasons for drop out. BMJ Open 9(6):e021959

  104. [113]

    Goodson N, Wicks P, Morgan J, Hashem L, Callinan S, Reites J. 2022. Opportunities and counterintuitive challenges for decentralized clinical trials to broaden participant inclusion. NPJ Digital Medicine 5:58

  105. [114]

    Thomas KA, Kidzi´ nski L. 2022. Artificial intelligence can improve patients’ experience in decentralized clinical trials. Nature Medicine 28(12):2462–2463

  106. [115]

    Zhou Q, Ratcliffe SJ, Grady C, Wang T, Mao JJ, Ulrich CM. 2019. Cancer Clinical Trial Patient-Participants’ Perceptions about Provider Communication and Dropout Intentions. AJOB empirical bioethics 10(3):190–200

  107. [116]

    Dennst¨ adt F, Zink J, Putora PM, Hastings J, Cihoric N. 2024. Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain. Systematic Reviews 13(1):158

  108. [117]

    Chen RJ, Lu MY, Chen TY, Williamson DFK, Mahmood F. 2021. Synthetic data in machine learning for medicine and healthcare. Nature Biomedical Engineering 5(6):493–497

  109. [118]

    Khosravi B, Li F, Dapamede T, Rouzrokh P, Gamble CU, et al. 2024. Synthetically enhanced: unveiling synthetic data’s potential in medical imaging research. eBioMedicine 104:105174

  110. [119]

    Das HP, Tran R, Singh J, Yue X, Tison G, et al. 2021. Conditional Synthetic Data Generation for Robust Machine Learning Applications with Limited Pandemic Data

  111. [120]

    Ive J, Viani N, Kam J, Yin L, Verma S, et al. 2020. Generation and evaluation of artificial mental health records for Natural Language Processing. npj Digital Medicine 3(1):1–9

  112. [121]

    Pierson E, Shanmugam D, Movva R, Kleinberg J, Agrawal M, et al. 2024. Use large language models to promote health equity. arXiv preprint arXiv:2312.14804

  113. [122]

    Sun Y, Zhu C, Zheng S, Zhang K, Sun L, et al. 2024. PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology

  114. [123]

    Lu MY, Chen B, Williamson DFK, Chen RJ, Zhao M, et al. 2024. A Multimodal Generative AI Copilot for Human Pathology. Nature :1–3

  115. [124]

    Huang Z, Bianchi F, Yuksekgonul M, Montine TJ, Zou J. 2023. A visual-language foundation model for pathology image analysis using medical Twitter. Nature Medicine 29(9):2307–2316

  116. [125]

    Wang H, Fu T, Du Y, Gao W, Huang K, et al. 2023. Scientific discovery in the age of artificial intelligence. Nature 620(7972):47–60 www.annualreviews.org • Generative AI in Medicine 23

  117. [126]

    Wong HE, Rakic M, Guttag J, Dalca A V. 2024. ScribblePrompt: Fast and Flexible Interactive Segmentation for Any Biomedical Image

  118. [127]

    Zhou Y, Liu H, Srivastava T, Mei H, Tan C. 2024. Hypothesis Generation with Large Language Models

  119. [128]

    Zhong R, Zhang P, Li S, Ahn J, Klein D, Steinhardt J. 2023. Goal Driven Discovery of Distributional Differences via Language Descriptions

  120. [129]

    Pham CM, Hoyle A, Sun S, Resnik P, Iyyer M. 2024. TopicGPT: A Prompt-based Topic Modeling Framework

  121. [130]

    Kamienny PA, d’Ascoli S, Lample G, Charton F. 2022. End-to-end symbolic regression with transformers. Advances in Neural Information Processing Systems 35:10269–10281

  122. [131]

    Tayebi Arasteh S, Han T, Lotfinia M, Kuhl C, Kather JN, et al. 2024. Large language models streamline automated machine learning for clinical studies. Nature Communications 15(1):1603

  123. [132]

    Mozannar H, Chen V, Alsobay M, Das S, Zhao S, et al. 2024. The RealHumanEval: Evaluating Large Language Models’ Abilities to Support Programmers

  124. [133]

    Biri SK, Kumar S, Panigrahi M, Mondal S, Behera JK, Mondal H. 2024. Assessing the Utiliza- tion of Large Language Models in Medical Education: Insights From Undergraduate Medical Students. Cureus 15(10):e47468

  125. [134]

    Grigorian A, Shipley J, Nahmias J, Nguyen N, Schwed AC, et al. 2023. Implications of Using Chatbots for Future Surgical Education. JAMA Surgery 158(11):1220–1222

  126. [135]

    Lee CR, Gilliland KO, Beck Dallaghan GL, Tolleson-Rinehart S. 2022. Race, ethnicity, and gender representation in clinical case vignettes: a 20-year comparison between two institutions. BMC Medical Education 22(1):585

  127. [136]

    Benoit JR. 2023. Chatgpt for clinical vignette generation, revision, and evaluation. medRxiv :2023–02

  128. [137]

    Tejani AS, Elhalawani H, Moy L, Kohli M, Kahn CE. 2023. Artificial Intelligence and Radi- ology Education. Radiology: Artificial Intelligence 5(1):e220084

  129. [138]

    Holderried F, Stegemann-Philipps C, Herschbach L, Moldt JA, Nevins A, et al. 2024. A Gen- erative Pretrained Transformer (GPT)-Powered Chatbot as a Simulated Patient to Practice History Taking: Prospective, Mixed Methods Study. JMIR medical education 10:e53961

  130. [139]

    Dai W, Lin J, Jin H, Li T, Tsai YS, et al. 2023. Can Large Language Models Provide Feed- back to Students? A Case Study on ChatGPT . In 2023 IEEE International Conference on Advanced Learning Technologies (ICALT), pp. 323–325

  131. [140]

    Belmar F, Gaete MI, Escalona G, Carnier M, Dur´ an V, et al. 2023. Artificial intelligence in laparoscopic simulation: a promising future for large-scale automated evaluations. Surgical Endoscopy 37(6):4942–4946

  132. [141]

    Zhao Y, Wang Y, Zhang J, Liu X, Li Y, et al. 2022. Surgical GAN: Towards real-time path planning for passive flexible tools in endovascular surgeries. Neurocomputing 500:567–580

  133. [142]

    Kirby MD. 1983. Informed consent: what does it mean? Journal of medical ethics 9(2):69–75

  134. [143]

    Riddick F A. 2003. The code of medical ethics of the american medical association

  135. [144]

    Del Carmen MG, Joffe S. 2005. Informed consent for medical treatment and research: a review. The oncologist 10(8):636–641

  136. [145]

    Pierson L, Pierson E. 2022. Patients cannot consent to care unless they know how much it costs

  137. [146]

    Astromsk˙ e K, Peiˇ cius E, Astromskis P. 2021. Ethical and legal challenges of informed consent applying artificial intelligence in medical diagnostic consultations.AI & SOCIETY 36:509–520

  138. [147]

    Wilcox L, Brewer R, Diaz F. 2023. Ai consent futures: A case study on voice data collection with clinicians. Proceedings of the ACM on Human-Computer Interaction 7(CSCW2):1–30

  139. [148]

    Garcia Valencia OA, Suppadungsuk S, Thongprayoon C, Miao J, Tangpanithandee S, et al

  140. [149]

    Decker H, Trang K, Ramirez J, Colley A, Pierce L, et al. 2023. Large language model- based chatbot vs surgeon-generated informed consent documentation for common procedures.JAMA Network Open 6(10):e2336997–e2336997

  141. [150]

    Burks AC, Keim-Malpass J. 2019. Health literacy and informed consent for clinical trials: a systematic review and implications for nurses. Nursing: Research and Reviews :31–40

  142. [151]

    Simon C, Zyzanski SJ, Eder M, Raiz P, Kodish ED, Siminoff LA. 2003. Groups potentially at risk for making poorly informed decisions about entry into clinical trials for childhood cancer. Journal of Clinical Oncology 21(11):2173–2178

  143. [152]

    Raimann FJ, Neef V, Hennighausen MC, Zacharowski K, Flinspach AN. 2024. Evaluation of ai chatbots for the creation of patient-informed consent sheets.Machine Learning and Knowledge Extraction 6(2):1145–1153

  144. [153]

    Chen Y, Esmaeilzadeh P. 2024. Generative ai in medical practice: in-depth exploration of privacy and security challenges. Journal of Medical Internet Research 26:e53008

  145. [154]

    Bai X, Wang H, Ma L, Xu Y, Gan J, et al. 2021. Advancing covid-19 diagnosis with privacy- preserving collaboration in artificial intelligence. Nature Machine Intelligence 3(12):1081–1089

  146. [155]

    Ali M, Naeem F, Tariq M, Kaddoum G. 2022. Federated learning for privacy preservation in smart healthcare systems: A comprehensive survey. IEEE journal of biomedical and health informatics 27(2):778–789

  147. [156]

    Xu J, Glicksberg BS, Su C, Walker P, Bian J, Wang F. 2021. Federated learning for healthcare informatics. Journal of healthcare informatics research 5:1–19

  148. [157]

    Geiping J, Bauermeister H, Dr¨ oge H, Moeller M. 2020. Inverting Gradients – How easy is it to break privacy in federated learning?

  149. [158]

    So J, Ali RE, Guler B, Jiao J, Avestimehr S. 2023. Securing Secure Aggregation: Mitigating Multi-Round Privacy Leakage in Federated Learning

  150. [159]

    Huang Y, Gupta S, Song Z, Li K, Arora S. 2021. Evaluating Gradient Inversion Attacks and Defenses in Federated Learning

  151. [160]

    Mullainathan S, Obermeyer Z. 2022. Solving medicine’s data bottleneck: Nightingale open science. Nature Medicine 28(5):897–899

  152. [161]

    Johnson AE, Bulgarelli L, Shen L, Gayles A, Shammout A, et al. 2023. Mimic-iv, a freely accessible electronic health record dataset. Scientific data 10(1):1

  153. [162]

    Barrett C, Boyd B, Bursztein E, Carlini N, Chen B, et al. 2023. Identifying and mitigating the security risks of generative ai. Foundations and Trends® in Privacy and Security 6(1):1–52

  154. [163]

    El-Mhamdi EM, Farhadkhani S, Guerraoui R, Gupta N, Hoang LN, et al. 2022. On the im- possible safety of large ai models. arXiv preprint arXiv:2209.15259

  155. [164]

    Carlini N, Ippolito D, Jagielski M, Lee K, Tramer F, Zhang C. 2022. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646

  156. [165]

    Huang J, Shao H, Chang KCC. 2022. Are large pre-trained language models leaking your personal information? arXiv preprint arXiv:2205.12628

  157. [166]

    Carlini N, Tramer F, Wallace E, Jagielski M, Herbert-Voss A, et al. 2021. Extracting training data from large language models . In 30th USENIX Security Symposium (USENIX Security 21), pp. 2633–2650

  158. [167]

    Choi E, Biswal S, Malin B, Duke J, Stewart WF, Sun J. 2018. Generating Multi-label Discrete Patient Records using Generative Adversarial Networks

  159. [168]

    Ghosheh GO, Li J, Zhu T. 2024. A survey of generative adversarial networks for synthesizing structured electronic health records. ACM Computing Surveys 56(6):1–34

  160. [169]

    Loong B, Zaslavsky AM, He Y, Harrington DP. 2013. Disclosure Control using Partially Syn- thetic Data for Large-Scale Health Surveys, with Applications to CanCORS. Statistics in medicine 32(24):4139–4161

  161. [170]

    Bommasani R, Klyman K, Longpre S, Kapoor S, Maslej N, et al. 2023. The foundation model transparency index. arXiv preprint arXiv:2310.12941

  162. [171]

    Bommasani R, et al. 2024. The foundation model transparency index v1.1 may 2024. Stanford www.annualreviews.org • Generative AI in Medicine 25 CRFM

  163. [172]

    Winkler JK, Fink C, Toberer F, Enk A, Deinlein T, et al. 2019. Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolu- tional neural network for melanoma recognition. JAMA dermatology 155(10):1135–1141

  164. [173]

    2020.Hidden stratification causes clinically meaningful failures in machine learning for medical imaging

    Oakden-Rayner L, Dunnmon J, Carneiro G, R´ e C. 2020.Hidden stratification causes clinically meaningful failures in machine learning for medical imaging . In Proceedings of the ACM conference on health, inference, and learning , pp. 151–159

  165. [174]

    Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. 2018. Confounding variables can degrade generalization performance of radiological deep learning models. arXiv preprint arXiv:1807.00431

  166. [175]

    Gilpin LH, Bau D, Yuan BZ, Bajwa A, Specter M, Kagal L. 2018. Explaining explanations: An overview of interpretability of machine learning . In 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA) , pp. 80–89. IEEE

  167. [176]

    Stiglic G, Kocbek P, Fijacko N, Zitnik M, Verbert K, Cilar L. 2020. Interpretability of machine learning-based prediction models in healthcare. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 10(5):e1379

  168. [177]

    Ghassemi M, Oakden-Rayner L, Beam AL. 2021. The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health 3(11):e745–e750

  169. [178]

    Bilodeau B, Jaques N, Koh PW, Kim B. 2024. Impossibility theorems for feature attribution. Proceedings of the National Academy of Sciences 121(2):e2304406120

  170. [179]

    Zhao H, Yang F, Lakkaraju H, Du M. 2024. Opening the black box of large language models: Two views on holistic interpretability. arXiv preprint arXiv:2402.10688

  171. [180]

    Singh C, Inala JP, Galley M, Caruana R, Gao J. 2024. Rethinking interpretability in the era of large language models. arXiv preprint arXiv:2402.01761

  172. [181]

    Agarwal C, Tanneru SH, Lakkaraju H. 2024. Faithfulness vs. plausibility: On the (un) relia- bility of explanations from large language models. arXiv preprint arXiv:2402.04614

  173. [182]

    U N, M K, J K. 2023. Vision-Language Transformer for Interpretable Pathology Visual Ques- tion Answering. IEEE journal of biomedical and health informatics 27(4)

  174. [183]

    Kim C, Gadgil SU, DeGrave AJ, Omiye JA, Cai ZR, et al. 2024. Transparent medical image ai via an image–text foundation model grounded in medical literature. Nature Medicine :1–12

  175. [184]

    Chen S, Kann BH, Foote MB, Aerts HJWL, Savova GK, et al. 2023. Use of Artificial Intelli- gence Chatbots for Cancer Treatment Information. JAMA Oncology

  176. [185]

    Ji Z, Lee N, Frieske R, Yu T, Su D, et al. 2023. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12):248:1–248:38

  177. [186]

    Lee P, Bubeck S, Petro J. 2023. Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine. New England Journal of Medicine 388(13):1233–1239

  178. [187]

    Pan A, Musheyev D, Bockelman D, Loeb S, Kabarriti AE. 2023. Assessment of Artificial Intelligence Chatbot Responses to Top Searched Queries About Cancer. JAMA Oncology

  179. [188]

    Shuster K, Poff S, Chen M, Kiela D, Weston J. 2021. Retrieval Augmentation Reduces Hallu- cination in Conversation

  180. [189]

    Agrawal M, Hegselmann S, Lang H, Kim Y, Sontag D. 2022. Large language models are few-shot clinical information extractors . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pp. 1998–2022

  181. [190]

    Gero Z, Singh C, Cheng H, Naumann T, Galley M, et al. 2023. Self-Verification Improves Few-Shot Clinical Information Extraction. ArXiv:2306.00024 [cs]

  182. [191]

    Singhal K, Tu T, Gottweis J, Sayres R, Wulczyn E, et al. 2023. Towards Expert-Level Medical Question Answering with Large Language Models. ArXiv:2305.09617 [cs]

  183. [192]

    Mesk´ o B, Topol EJ. 2023. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digital Medicine 6(1):1–6Number: 1 Publisher: Nature Publishing Group

  184. [193]

    Jakob Nielsen. 2023. AI: First New UI Paradigm in 60 Years. Nielsen Norman Group 26 Shanmugam et al

  185. [194]

    Sai S, Gaur A, Sai R, Chamola V, Guizani M, Rodrigues JJ. 2024. Generative AI for trans- formative healthcare: A comprehensive study of emerging models, applications, case studies and limitations. IEEE Access

  186. [195]

    Mulia AP, Piri PR, Tho C. 2023. Usability analysis of text generation by chatgpt openai using system usability scale method. Procedia Computer Science 227:381–388

  187. [196]

    Giunti G, Doherty CP. 2024. Cocreating an automated mhealth apps systematic review process with generative ai: Design science research approach. JMIR Medical Education 10:e48949

  188. [197]

    Tankelevitch L, Kewenig V, Simkute A, Scott AE, Sarkar A, et al. 2024. The metacognitive demands and opportunities of generative AI . In Proceedings of the CHI Conference on Human Factors in Computing Systems , pp. 1–24

  189. [198]

    Dang H, Mecke L, Lehmann F, Goller S, Buschek D. 2022. How to prompt? Opportunities and challenges of zero-and few-shot learning for human-AI interaction in creative applications of generative models. arXiv preprint arXiv:2209.01390

  190. [199]

    Sun J, Liao QV, Muller M, Agarwal M, Houde S, et al. 2022. Investigating explainability of generative AI for code through scenario-based design. In Proceedings of the 27th International Conference on Intelligent User Interfaces , pp. 212–228

  191. [200]

    Subramonyam H, Pea R, Pondoc C, Agrawala M, Seifert C. 2024. Bridging the Gulf of En- visioning: Cognitive Challenges in Prompt Based Interactions with LLMs . In Proceedings of the CHI Conference on Human Factors in Computing Systems , pp. 1–19

  192. [201]

    Zamfirescu-Pereira J, Wong RY, Hartmann B, Yang Q. 2023. Why Johnny can ’t prompt: how non-AI experts try (and fail) to design LLM prompts . In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , pp. 1–21

  193. [202]

    Abbasian M, Khatibi E, Azimi I, Oniani D, Shakeri Hossein Abad Z, et al. 2024. Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative ai. NPJ Digital Medicine 7(1):82

  194. [203]

    Vasconcelos H, J¨ orke M, Grunde-McLaughlin M, Gerstenberg T, Bernstein MS, Krishna R

  195. [204]

    Kostick-Quenet KM, Gerke S. 2022. AI in the hands of imperfect users. npj Digital Medicine 5(1):197

  196. [205]

    Proceedings of the ACM on Human-Computer Interaction 7(CSCW1):1–38

    Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction 7(CSCW1):1–38

  197. [206]

    Jacobs M, Pradier MF, McCoy Jr TH, Perlis RH, Doshi-Velez F, Gajos KZ. 2021. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry 11(1):108

  198. [207]

    Wysocki O, Davies JK, Vigo M, Armstrong AC, Landers D, et al. 2023. Assessing the com- munication gap between AI models and healthcare professionals: Explainability, utility and trust in AI-driven clinical decision-making. Artificial Intelligence 316:103839

  199. [208]

    Schaefer KE, Chen JY, Szalma JL, Hancock PA. 2016. A meta-analysis of factors influencing the development of trust in automation: Implications for understanding autonomy in future systems. Human factors 58(3):377–400

  200. [209]

    2024.Design Principles for Generative AI Applications

    Weisz JD, He J, Muller M, Hoefer G, Miles R, Geyer W. 2024.Design Principles for Generative AI Applications . In Proceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1–22

  201. [210]

    Greenes RA, Bates DW, Kawamoto K, Middleton B, Osheroff J, Shahar Y. 2018. Clinical decision support models and frameworks: seeking to address research issues underlying imple- mentation successes and failures. Journal of biomedical informatics 78:134–143

  202. [211]

    Mertz L. 2015. From annoying to appreciated: Turning clinical decision support systems into a medical professional’s best friend. IEEE pulse 6(5):4–9

  203. [212]

    Khan S, Richardson S, Liu A, Mechery V, McCullagh L, et al. 2019. Improving provider www.annualreviews.org • Generative AI in Medicine 27 adoption with adaptive clinical decision support surveillance: an observational study. JMIR human factors 6(1):e10245

  204. [213]

    Wright A, Hickman TTT, McEvoy D, Aaron S, Ai A, et al. 2016. Analysis of clinical decision support system malfunctions: a case series and survey. Journal of the American Medical Informatics Association 23(6):1068–1076

  205. [214]

    Gaube S, Suresh H, Raue M, Merritt A, Berkowitz SJ, et al. 2021. Do as AI say: susceptibility in deployment of clinical decision-aids. NPJ digital medicine 4(1):31

  206. [215]

    Mann D, Hess R, McGinn T, Richardson S, Jones S, et al. 2020. Impact of clinical deci- sion support on antibiotic prescribing for acute respiratory infections: a cluster randomized implementation trial. Journal of General Internal Medicine 35:788–795

  207. [216]

    Pradier M, Lam B, Ahn AC, et al

    Jacobs M, He J, F. Pradier M, Lam B, Ahn AC, et al. 2021. Designing AI for trust and collaboration in time-constrained medical decisions: a sociotechnical lens . In Proceedings of the 2021 chi conference on human factors in computing systems , pp. 1–14

  208. [217]

    Gaube S, Suresh H, Raue M, Lermer E, Koch TK, et al. 2023. Non-task expert physicians benefit from correct explainable ai advice when reviewing x-rays. Scientific reports 13(1):1383

  209. [218]

    Cabrera AA, Perer A, Hong JI. 2023. Improving Human-AI Collaboration With Descriptions of AI Behavior. Proc. ACM Hum.-Comput. Interact. 7(CSCW1):136:1–136:21

  210. [219]

    Henry KE, Adams R, Parent C, Soleimani H, Sridharan A, et al. 2022. Factors driving provider adoption of the TREWS machine learning-based early warning system and its effects on sepsis treatment timing. Nature medicine 28(7):1447–1454

  211. [220]

    Gen AI Saves Nurses Time by Drafting Responses to Patient Messages — epicshare.org

    2024. Gen AI Saves Nurses Time by Drafting Responses to Patient Messages — epicshare.org. https://www.epicshare.org/share-and-learn/mayo-ai-message-responses

  212. [221]

    Bansal G, Nushi B, Kamar E, Lasecki WS, Weld DS, Horvitz E. 2019. Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 7:2–11

  213. [222]

    Gichoya JW, Banerjee I, Bhimireddy AR, Burns JL, Celi LA, et al. 2022. Ai recognition of patient race in medical imaging: a modelling study. The Lancet Digital Health 4(6):e406–e414

  214. [223]

    Obermeyer Z, Powers B, Vogeli C, Mullainathan S. 2019. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464):447–453

  215. [224]

    Gervasi SS, Chen IY, Smith-McLallen A, Sontag D, Obermeyer Z, et al. 2022. The potential for bias in machine learning and opportunities for health insurers to address it. Health Affairs 41(2):212–218

  216. [225]

    Daneshjou R, Vodrahalli K, Novoa RA, Jenkins M, Liang W, et al. 2022. Disparities in derma- tology ai performance on a diverse, curated clinical image set.Science advances 8(31):eabq6147

  217. [226]

    Chen IY, Pierson E, Rose S, Joshi S, Ferryman K, Ghassemi M. 2021. Ethical machine learning in healthcare. Annual review of biomedical data science 4:123–144

  218. [227]

    Zink A, Obermeyer Z, Pierson E. 2024. Race adjustments in clinical algorithms can help correct for racial disparities in data quality. PNAS

  219. [228]

    Wiens J, Saria S, Sendak M, Ghassemi M, Liu VX, et al. 2019. Do no harm: a roadmap for responsible machine learning for health care. Nature medicine 25(9):1337–1340

  220. [229]

    Pierson E. 2020. Assessing racial inequality in covid-19 testing with bayesian threshold tests. Extended abstract, NeurIPS ML4H

  221. [230]

    Seyyed-Kalantari L, Zhang H, McDermott MB, Chen IY, Ghassemi M. 2021. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nature medicine 27(12):2176–2182

  222. [231]

    Pierson E. 2024. Accuracy and equity in clinical risk prediction. The New England Journal of Medicine 390(2):100–102

  223. [232]

    Zou J, Gichoya JW, Ho DE, Obermeyer Z. 2023. Implications of predicting race variables from medical images. Science 381(6654):149–150

  224. [233]

    Movva R, Shanmugam D, Hou K, Pathak P, Guttag J, et al. 2023. Coarse race data conceals disparities in clinical risk score performance. In Machine Learning for Healthcare Conference, pp. 443–472. PMLR

  225. [234]

    Ferryman K, Mackintosh M, Ghassemi M. 2023. Considering biased data as informative arti- facts in AI-assisted health care. New England Journal of Medicine 389(9):833–838

  226. [235]

    Balachandar S, Garg N, Pierson E. 2024. Domain constraints improve risk prediction when outcome data is missing. ICLR 28 Shanmugam et al

  227. [236]

    Zink A, Rose S. 2020. Fair regression for health care spending. Biometrics 76(3):973–982

  228. [237]

    Shanmugam D, Hou K, Pierson E. 2024. Quantifying disparities in intimate partner violence: a machine learning method to correct for underreporting. npj Women ’s Health 2(1):15

  229. [238]

    Mullainathan S, Obermeyer Z. 2021. On the inequity of predicting A while hoping for B . In AEA Papers and Proceedings , vol. 111, pp. 37–42. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203

  230. [239]

    Vyas DA, Eisenstein LG, Jones DS. 2020. Hidden in plain sight—reconsidering the use of race correction in clinical algorithms

  231. [240]

    Obermeyer Z, Nissan R, Stern M, Eaneff S, Bembeneck EJ, Mullainathan S. 2021. Algorithmic bias playbook. Center for Applied AI at Chicago Booth

  232. [241]

    Pierson E, Cutler DM, Leskovec J, Mullainathan S, Obermeyer Z. 2021. An algorithmic ap- proach to reducing unexplained pain disparities in underserved populations. Nature Medicine 27(1):136–140

  233. [242]

    Diao JA, Shi I, Murthy VL, Buckley TA, Patel CJ, et al. 2024. Projected changes in statin and antihypertensive therapy eligibility with the aha prevent cardiovascular risk equations. JAMA 332(12):989–1000

  234. [243]

    Diao JA, He Y, Khazanchi R, Tiako MN, Witonsky JI, et al. 2024. Implications of race adjustment in lung-function equations. The New England journal of medicine 390(22):2083

  235. [244]

    Hofmann V, Kalluri PR, Jurafsky D, King S. 2024. AI generates covertly racist decisions about people based on their dialect. Nature 633(8028):147–154

  236. [245]

    Pfohl SR, Cole-Lewis H, Sayres R, Neal D, Asiedu M, et al. 2024. A toolbox for surfacing health equity harms and biases in large language models. Nature Medicine :1–11

  237. [246]

    Zack T, Lehman E, Suzgun M, Rodriguez JA, Celi LA, et al. 2024. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet Digital Health 6(1):e12–e22

  238. [247]

    Hoffman KM, Trawalter S, Axt JR, Oliver MN. 2016. Racial bias in pain assessment and treatment recommendations, and false beliefs about biological differences between blacks and whites. Proceedings of the National Academy of Sciences of the United States of America 113(16):4296–4301

  239. [248]

    Ripp K, Braun L. 2017. Race/Ethnicity in Medical Education: An Analysis of a Question Bank for Step 1 of the United States Medical Licensing Examination. Teaching and Learning in Medicine 29(2):115–122

  240. [249]

    Omiye JA, Lester J, Spichak S, Rotemberg V, Daneshjou R. 2023. Beyond the hype: large language models propagate race-based medicine. medRxiv :2023–07

  241. [250]

    Weidinger L, Uesato J, Rauh M, Griffin C, Huang PS, et al. 2022. Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 214–229

  242. [251]

    Vogels EA. 2023. A majority of Americans have heard of ChatGPT, but few have tried it themselves

  243. [252]

    Smith B, Magnani JW. 2019. New technologies, new disparities: The intersection of electronic health and digital health literacy. International Journal of Cardiology 292:280–282Publisher: Elsevier

  244. [253]

    Veinot TC, Mitchell H, Ancker JS. 2018. Good intentions are not enough: how informatics interventions can worsen inequality.Journal of the American Medical Informatics Association: JAMIA 25(8):1080–1088

  245. [254]

    Rodriguez JA, Alsentzer E, Bates DW. 2024. Leveraging large language models to foster equity www.annualreviews.org • Generative AI in Medicine 29 in healthcare. Journal of the American Medical Informatics Association :ocae055

  246. [255]

    Long D, Magerko B. 2020. What is AI Literacy? Competencies and Design Considerations . In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , CHI ’20, pp. 1–16. New York, NY, USA: Association for Computing Machinery

  247. [256]

    Jin D, Pan E, Oufattole N, Weng WH, Fang H, Szolovits P. 2020. What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. ArXiv:2009.13081 [cs]

  248. [257]

    Kim JY, Hasan A, Kellogg KC, Ratliff W, Murray SG, et al. 2024. Development and prelimi- nary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities. PLOS Digi...

  249. [258]

    Jin Q, Dhingra B, Liu Z, Cohen WW, Lu X. 2019. PubMedQA: A Dataset for Biomedical Research Question Answering. ArXiv:1909.06146 [cs, q-bio]

  250. [259]

    Pal A, Umapathi LK, Sankarasubbu M. 2022. MedMCQA : A Large-scale Multi-Subject Multi- Choice Dataset for Medical domain Question Answering. ArXiv:2203.14371 [cs]

  251. [260]

    Deas N, Grieser J, Kleiner S, Patton D, Turcan E, McKeown K. 2023. Evaluation of African American Language Bias in Natural Language Generation. ArXiv:2305.14291 [cs]

  252. [261]

    Lai VD, Ngo NT, Veyseh APB, Man H, Dernoncourt F, et al. 2023. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. ArXiv:2304.05613 [cs]

  253. [262]

    Abacha AB, Mrabet Y, Sharp M, Goodwin TR, Shooshan SE, Demner-Fushman D. 2019. Bridging the Gap Between Consumers’ Medication Questions and Trusted Answers. Studies in Health Technology and Informatics 264:25–29

  254. [263]

    Ghosh S, Caliskan A. 2023. ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Lan- guages. arXiv preprint arXiv:2305.10510

  255. [264]

    Yeo YH, Samaan JS, Ng WH, Ting PS, Trivedi H, et al. 2023. Assessing the performance of ChatGPT in answering questions regarding cirrhosis and hepatocellular carcinoma. medRxiv :2023–02

  256. [265]

    Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, et al. 2023. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine

  257. [266]

    Li SS, Balachandran V, Feng S, Ilgen J, Pierson E, et al. 2024. MEDIQ: Question-Asking LLMs for Adaptive and Reliable Clinical Reasoning

  258. [267]

    Ayers JW, Zhu Z, Poliak A, Leas EC, Dredze M, et al. 2023. Evaluating artificial intelligence responses to public health questions. JAMA Network Open 6(6):e2317517–e2317517

  259. [268]

    Cestonaro C, Delicati A, Marcante B, Caenazzo L, Tozzo P. 2023. Defining medical liability when artificial intelligence is applied on diagnostic algorithms: a systematic review. Frontiers in Medicine 10:1305756

  260. [269]

    Antoniak M, Naik A, Alvarado CS, Wang LL, Chen IY. 2024. NLP for Maternal Healthcare: Perspectives and Guiding Principles in the Age of LLMs . In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pp. 1446–1463

  261. [270]

    Vicente L, Matute H. 2023. Humans inherit artificial intelligence biases. Scientific Reports 13(1):15737

  262. [271]

    Bastani H, Bastani O, Sungu A, Ge H, Kabakcı O, Mariman R. 2024. Generative AI Can Harm Learning. Available at SSRN 4895486

  263. [273]

    Sung JJ, Poon NC. 2020. Artificial intelligence in gastroenterology: where are we heading? Frontiers of medicine 14(4):511–517 30 Shanmugam et al

  264. [2023]

    Journal of Personalized Medicine 13(9):1363 24 Shanmugam et al

    Ethical implications of chatbot utilization in nephrology. Journal of Personalized Medicine 13(9):1363 24 Shanmugam et al

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.