Pith. sign in

REVIEW 2 major objections 8 minor 78 references

Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action

T0 review · 2 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper argues that the total product life cycle framework cannot adequately regulate LLM-based medical devices.

desk verdict A solid regulatory perspective whose technical premise overreaches; the TPLC argument survives once 'non-deterministic' is qualified. read the letter →

arxiv 2502.07794 v1 pith:7UWAATM5 submitted 2025-01-27 cs.CY cs.AI

classification cs.CYcs.AI
keywords generativeAIlargelanguagemodelsmedicaldeviceregulationtotalproductlifecycleregulatorysandboxesadaptivehealthequityscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This perspective paper argues that the total product life cycle (TPLC) framework—the standard regulatory structure that follows a medical device from design through post-market monitoring—cannot adequately govern LLM-based medical devices. The authors identify three structural reasons: LLM outputs are non-deterministic even when temperature is set to zero; LLM functions are broad and generalist, making intended-use classification ambiguous; and LLM systems are complex composites whose behavior shifts with prompts, model versions, and third-party modifications. Because of this, premarket verification cannot certify safety at a single time point, and automated post-market monitoring cannot capture long-form outputs. The paper concludes that regulators should move toward adaptive policies, regulatory sandboxes, international harmonization, and supply-chain oversight, developed through global multidisciplinary regulatory science. A sympathetic reader would care because the current regulatory toolkit was built for deterministic, single-purpose software, and the safe deployment of generative AI in medicine depends on whether that toolkit is replaced in time.

What carries the argument

The central object is the total product life cycle (TPLC) framework, a regulatory structure that tracks a software medical device from planning and design through verification, deployment, and post-market monitoring. The paper uses TPLC as the baseline that LLM-based devices fail, and the load-bearing property is the assertion that LLM outputs are non-deterministic even when the temperature is set to zero. That property is what makes one-time verification insufficient and continuous monitoring necessary but hard. The proposed replacement mechanism is the regulatory sandbox—a constrained real-world environment in which developers and regulators test policies iteratively before full approval—paired with adaptive policies that tighten or loosen restrictions as real-world evidence accumulates. The machinery also includes international harmonization of evaluation metrics and data provenance standards as the coordinating layer.

What would settle it

Run the same medically relevant prompt through the same model hundreds of times with temperature set to zero and identical sampling configuration, and measure the distribution of outputs; if outputs are identical across runs and stable across controlled implementations, the premise of intrinsic non-determinism would be falsified, and the case for abandoning TPLC-style verification would weaken.

Watch

Extended reading notes

Core claim

The central claim is that TPLC, as currently practiced, is the wrong instrument for LLM-based medical devices. The paper's argument turns on three properties of LLMs: outputs are non-deterministic by nature even at zero temperature; functionality is not tied to a single intended use but spans summarization, diagnosis suggestions, and documentation; and integration is layered, so fine-tuning, retrieval-augmented generation, prompt variation, and base-model updates all change behavior after approval. Each of these properties breaks a different phase of the TPLC: verification and validation cannot be completed once, operation and monitoring cannot be automated for free-form outputs, and classification and enforcement cannot rely on intended-use definitions or predicate-based clearance. The paper's positive thesis is that regulatory science must innovate through adaptive regulation and regulatory sandboxes, with global harmonization of standards and deliberate attention to health equity, so that governance can be tested and revised in real-world settings rather than fixed at approval.

Load-bearing premise

The argument leans on the claim that LLM outputs are inherently non-deterministic even when the model temperature is set to zero, so that no fixed verification can pin down their behavior; this is asserted rather than measured.

Editorial extensions

If this is right

  • Regulators would need to treat LLM-based medical devices as a distinct regulatory category rather than fitting them into existing single-purpose software pathways.
  • Accelerated approval routes based on substantial equivalence to predicate devices would have to be revisited, because fine-tuning and retrieval-augmented generation can change a model's risk profile without a new submission.
  • Post-market surveillance of LLM tools would require human-in-the-loop review of long-form outputs and new incident-reporting mechanisms, since automated drift detection cannot capture hallucinations or redaction errors.
  • Global harmonization of evaluation metrics, dataset-provenance standards, and risk classification would become a prerequisite for managing cross-border deployment of LLM medical tools.
  • Regulatory sandboxes and adaptive policies would become standard tools for generating real-world evidence before and after approval.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If non-determinism is the true crux, then a direct measurement campaign—running identical prompts on identical models with temperature zero, fixed seeds, and controlled decoding—would either confirm or undercut the paper's central premise, which is asserted rather than demonstrated.
  • The same TPLC mismatch likely extends to high-stakes LLM deployment outside medicine, such as legal advice or financial decisions, where a single approval-time evaluation cannot guarantee behavior across future prompts and versions.
  • Regulatory sandboxes could be designed as comparative experiments across jurisdictions, publishing outcomes with common metrics so that different governance policies can be evaluated against each other.
  • The paper's observation that product and service regulation are converging in health suggests a broader shift: as LLM-based agents deliver services rather than fixed functions, product-centric regulation will increasingly need service-centric oversight.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 8 minor

Summary. This perspective argues that existing medical device regulatory frameworks, in particular the U.S. FDA's Total Product Life Cycle (TPLC) approach, are poorly suited to generative AI and large language models in healthcare. It identifies several classes of challenge: ambiguity in medical device definitions, lack of robust evaluation methods, difficulties in monitoring and enforcement, and ethical and equity concerns. It then proposes a global regulatory science agenda centered on adaptive regulatory approaches, regulatory sandboxes, supply-chain oversight, and international harmonization, with attention to health equity. The paper is a call to action rather than an empirical study, and its claims are supported primarily by cited literature and illustrative examples.

Significance. The essay is timely and well-referenced; it brings together recent evaluation studies (e.g., Hager et al.), reporting guidelines, and regulatory initiatives across the US, EU, UK, and Singapore. Its main value is as a synthesis and agenda-setting piece for regulatory science. It does not provide new measurements, formal models, or outcome data, and the central recommendation is programmatic. Nevertheless, the paper makes a useful contribution by framing the regulatory problem and identifying concrete gaps, such as predicate creep, data provenance, and supply-chain vulnerabilities, that are often discussed separately. If revised to qualify its stronger empirical claims, it would be a serviceable perspective for stimulating discussion at the intersection of AI, medicine, and regulation.

major comments (2)
  1. [Ambiguity in Medical Device Definitions] The sentence 'Unlike predictive models, LLM outputs are non-deterministic in nature, even when the model temperature is set to zero' is an overstatement that is not supported by the cited literature. With fixed weights, fixed infrastructure, and greedy decoding, LLM inference can be bit-identical across runs; observed API-level nondeterminism typically arises from batching, floating-point non-associativity, load balancing, or software updates. The paper should replace this with the weaker, defensible claim that LLM outputs are highly sensitive to prompts, contexts, model versions, and implementation choices, and support that claim with citations. This correction matters because the sentence is invoked to justify revising risk classification and control for LLM-based devices; however, the weaker claim is sufficient for the paper's TPLC critique, so the issue is locally fixable rather than fatal to the overall argument.
  2. [Applying Adaptive Regulatory Approaches] The proposal for regulatory sandboxes and adaptive policies would be strengthened by specifying measurable success criteria and data-collection mechanisms. As written, the sandbox discussion relies on the OECD characterization and examples such as MHRA's AI Airlock and Singapore's IMDA sandbox, but it does not say how a sandbox's outcome would be evaluated, how results would generalize across jurisdictions, or what would count as failure triggering withdrawal of a policy. For a 'global call for action,' the absence of an evaluation design is a nontrivial gap that leaves the central recommendation untestable.
minor comments (8)
  1. [Introduction] In the paragraph on TPLC, 'are adaption' should be 'are adopting.'
  2. [Table 1] In the Europe row, 'provsions' should be 'provisions.'
  3. [Advancing the Collective Goals of Heath Equity] The heading misspells 'Health' as 'Heath.'
  4. [Future Directions for Regulatory Science and Regulators] In the LMIC paragraph, 'the authors note raised the need' should be 'the authors noted the need.'
  5. [Beyond Medical Device Regulation: Responsible AI in Health Product Development] The phrase 'lends in to' should be 'lends itself to.'
  6. [Conclusion] The phrase 'challenges that that fall' should be 'challenges that fall.'
  7. [Challenges to Monitoring and Regulatory Enforcement] The phrase 'close to two-third' should be 'close to two-thirds.'
  8. [References] Reference formatting is inconsistent, with some entries using abbreviated author names (e.g., 'H-G, E., et al.' and 'Alan, B.'); a consistent author-year style would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this is a policy perspective with no fitted quantities or derived predictions, and its central claims rest on external evidence rather than on self-citation.

full rationale

This manuscript is a perspective/commentary, not a quantitative derivation. The strongest claim—that GenAI/LLM non-determinism, broad functionality, and complex integration challenge the total product life cycle (TPLC) approach—is argued through descriptive reasoning and citations to external sources such as Hager et al. (reference 7), FDA documents (reference 3), and Warraich et al. (reference 4). No equations are derived, no parameters are fitted, and no quantity is predicted from an input; therefore the main circularity failure modes (self-definitional steps, fitted inputs called predictions, renaming known results) do not apply. The paper does contain multiple self-citations by the author group (e.g., references 5, 23, 30, 56, 64, 65), but these support background claims about LLM ethics, clinical decision support, adverse event scoping, and reporting checklists; they are not the sole or load-bearing justification for the paper's central recommendation of global regulatory science collaboration. The sentence in the 'Ambiguity in Medical Device Definitions' section stating that 'LLM outputs are non-deterministic in nature, even when the model temperature is set to zero' is an empirical assertion offered without citation or measurement, but an unsupported or overbroad claim is a correctness/evidence concern, not a circularity concern: the paper does not derive this assertion from its own conclusion, nor does it define the conclusion in terms of the assertion. Because the paper is self-contained as a policy argument and does not reduce to its own inputs, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper adds no fitted numbers or derivations. It relies on several background assumptions about LLM behavior, regulatory classification, and the value of adaptive regulation, and these are asserted with citations rather than demonstrated. No invented physical or conceptual entities with independent falsifiable handles are introduced.

assumptions (3)
  • domain assumption LLM outputs are non-deterministic even when the temperature is set to zero.
    Invoked in Section 4 to justify why TPLC-style verification and monitoring cannot be applied; stated as fact without supporting measurement.
  • domain assumption Clinical documentation scribes that summarize physician-patient interactions influence clinical decisions and should therefore fall under medical device definitions.
    Used in the 'Ambiguity in Medical Device Definitions' section and Table 1 to argue that some marketed scribes should have required approval; the inference from summarizing to influencing decisions is asserted.
  • domain assumption Adaptive regulatory approaches and regulatory sandboxes improve safety and innovation.
    Advocated in 'Applying Adaptive Regulatory Approaches'; cited sources describe these tools conceptually, but no controlled evidence links them to better patient outcomes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action." pith.science (2026). https://pith.science/paper/7UWAATM5

@misc{pith2026250207794,
  author       = {Pith},
  title        = {Pith review of: Regulatory Science Innovation for Generative AI and Large Language Models in Health and Medicine: A Global Call for Action},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7UWAATM5}},
  note         = {Machine review of arXiv:2502.07794}
}
read the original abstract

The integration of generative AI (GenAI) and large language models (LLMs) in healthcare presents both unprecedented opportunities and challenges, necessitating innovative regulatory approaches. GenAI and LLMs offer broad applications, from automating clinical workflows to personalizing diagnostics. However, the non-deterministic outputs, broad functionalities and complex integration of GenAI and LLMs challenge existing medical device regulatory frameworks, including the total product life cycle (TPLC) approach. Here we discuss the constraints of the TPLC approach to GenAI and LLM-based medical device regulation, and advocate for global collaboration in regulatory science research. This serves as the foundation for developing innovative approaches including adaptive policies and regulatory sandboxes, to test and refine governance in real-world settings. International harmonization, as seen with the International Medical Device Regulators Forum, is essential to manage implications of LLM on global health, including risks of widening health inequities driven by inherent model biases. By engaging multidisciplinary expertise, prioritizing iterative, data-driven approaches, and focusing on the needs of diverse populations, global regulatory science research enables the responsible and equitable advancement of LLM innovations in healthcare.

Figures

Figures reproduced from arXiv: 2502.07794 by the authors.

Figure 1
Figure 1. An illustration of the different phases of TPLC for LLM-based medical devices, and the unique considerations and regulatory challenges at each phase. Planning and Design Data Collection and Management Model Building and Tuning Verification and Validation Model Deployment Operation and Monitoring Real-World Performance Evaluations LLM Total Product Life Cycle Planning and Design •Broad (vs narrowly defined) problem s… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 76 canonical work pages

  1. [1]

    & Hernandez-Boussard, T

    Ng, M.Y., Kapur, S., Blizinsky, K.D. & Hernandez-Boussard, T. The AI life cycle: a holistic approach to creating ethical AI for health decisions. Nature medicine 28(2022)

  2. [2]

    Real-World and Regulatory Perspectives of Artificial Intelligence in Cardiovascular Imaging

    Wellnhofer, E. Real-World and Regulatory Perspectives of Artificial Intelligence in Cardiovascular Imaging. Frontiers in cardiovascular medicine 9(2022)

  3. [3]

    Total Product Lifecycle Considerations for Generative AI- Enabled Devices

    Food and Drug Administration, U.S.A. Total Product Lifecycle Considerations for Generative AI- Enabled Devices. in Executive Summary for the Digital Health Advisory Committee Retrieved from: https://www.fda.gov/media/182871/download, Nov 2024

  4. [4]

    FDA Perspective on the Regulation of Artificial Intelligence in Health Care and Biomedicine

    Warraich, H.J., Tazbaz T., Califf Robert M. FDA Perspective on the Regulation of Artificial Intelligence in Health Care and Biomedicine. JAMA (2024)

  5. [5]

    Ethical and regulatory challenges of large language models in medicine

    Ong, J.C.L., et al. Ethical and regulatory challenges of large language models in medicine. The Lancet Digital health (2024)

  6. [6]

    & Vayena, E

    Aboy, M., Minssen, T. & Vayena, E. Navigating the EU AI Act: implications for regulated digital medical products. npj Digit. Med. 7, 237 (2024)

  7. [7]

    Hager, P., Jungmann, F., Holland, R. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med 30, 2613–2622 (2024) 14

  8. [8]

    Pfohl, S.R., Cole-Lewis, H., Sayres, R. et al. A toolbox for surfacing health equity harms and biases in large language models. Nat Med 30, 3590–3600 (2024)

Show all 78 references
  1. [9]

    Medical large language models are vulnerable to data-poisoning attacks

    Alber, D.A., et al. Medical large language models are vulnerable to data-poisoning attacks. Nature Medicine, 1-9 (2025)

  2. [10]

    Policing the boundary between responsible and irresponsible placing on the market of LLM health applications, Mayo Clinic Proceedings Digital Health, 2025 (in press)

    Freyer, Wiest, Gilbert. Policing the boundary between responsible and irresponsible placing on the market of LLM health applications, Mayo Clinic Proceedings Digital Health, 2025 (in press)

  3. [11]

    Longpre, S., Mahari, R., Chen, A. et al. A large-scale audit of dataset licensing and attribution in AI. Nat Mach Intell 6, 975–987 (2024)

  4. [12]

    Algorithmovigilance-Advancing Methods to Analyze and Monitor Artificial Intelligence- Driven Health Care for Effectiveness and Equity

    Peter J.E. Algorithmovigilance-Advancing Methods to Analyze and Monitor Artificial Intelligence- Driven Health Care for Effectiveness and Equity. JAMA network open 4(2021)

  5. [13]

    & Philippe, R

    Alan, B., Mehdi, B., Theodoros, E. & Philippe, R. Algorithmovigilance, lessons from pharmacovigilance. NPJ digital medicine 7(2024)

  6. [14]

    Off-label Medication Use: A Double-edged Sword

    Vandana, A. Off-label Medication Use: A Double-edged Sword. Indian journal of critical care medicine : peer-reviewed, official publication of Indian Society of Critical Care Medicine 25(2021)

  7. [15]

    challenging

    Caterina, P., et al. Limitations and obstacles of the spontaneous adverse drugs reactions reporting: Two "challenging" case reports. Journal of pharmacology & pharmacotherapeutics 4(2013)

  8. [16]

    & AC, van Grootheest

    L, Harmark. & AC, van Grootheest. Pharmacovigilance: methods, recent developments and future perspectives. European journal of clinical pharmacology 64(2008)

  9. [17]

    Pharmacovigilance: Importance, concepts, and processes

    Atul, K. Pharmacovigilance: Importance, concepts, and processes. American journal of health- system pharmacy : AJHP : official journal of the American Society of Health-System Pharmacists 74(2017)

  10. [18]

    & Kerstin N, V

    Urs J, M., Paola, D. & Kerstin N, V. Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015-20): a comparative analysis. The Lancet. Digital health 3(2021)

  11. [19]

    & Sandra, R

    Charlotte, L. & Sandra, R. Identification of predicate creep under the 510(k) process: A case study of a robotic surgical device. PloS one 18(2023)

  12. [20]

    & Kerstin N, V

    Urs J, M., Christain, B. & Kerstin N, V. FDA-cleared artificial intelligence and machine learning- based medical devices and their 510(k) predicate networks. The Lancet. Digital health 5(2023)

  13. [21]

    & Harlan M, K

    Kushal T, K., Sanket S, D., Cesar, C., Joseph S, R. & Harlan M, K. Use of Recalled Devices in New Device Authorizations Under the US Food and Drug Administration's 510(k) Pathway and Risk of Subsequent Recalls. JAMA 329(2023)

  14. [22]

    Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report

    Ke, Y., et al. Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report. Accessed in Arxiv: https://doi.org/10.48550/arXiv.2402.01733 (2024)

  15. [23]

    Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties

    Ong, J.C.L., et al. Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties. Accessed in Arxiv: https://doi.org/10.48550/arXiv.2402.01741 (2024)

  16. [24]

    Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs

    Li, W., et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. NPJ digital medicine 7(2024)

  17. [25]

    Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial

    Bertalan, M. Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial. Journal of medical Internet research 25(2023)

  18. [26]

    Closing the gap between open source and commercial large language models for medical evidence summarization

    Gongbo, Z., et al. Closing the gap between open source and commercial large language models for medical evidence summarization. NPJ digital medicine 7(2024)

  19. [27]

    Benchmarking Open-Source Large Language Models, GPT-4 and Claude 2 on Multiple-Choice Questions in Nephrology

    Wu, S., et al. Benchmarking Open-Source Large Language Models, GPT-4 and Claude 2 on Multiple-Choice Questions in Nephrology. (2024)

  20. [28]

    DeepLearning.AI

    Andrew Ng. DeepLearning.AI. Falling LLM Token Prices and What They Mean for AI Companies. Accessed at: https://www.deeplearning.ai/the-batch/falling-llm-token-prices-and-what-they-mean- for-ai-companies/ (2024)

  21. [29]

    A future role for health applications of large language models depends on regulators enforcing safety standards

    Freyer, O., et al. A future role for health applications of large language models depends on regulators enforcing safety standards. The Lancet Digital Health 6(2024)

  22. [30]

    Medical Ethics of Large Language Models in Medicine

    Ong, J.C.L., et al. Medical Ethics of Large Language Models in Medicine. NEJM AI 1(June 17 2024)

  23. [31]

    & Daneshjou, R

    Omiye, J.A., Lester, J.C., Spichak, S., Rotemberg, V. & Daneshjou, R. Large language models propagate race-based medicine. npj Digital Medicine 6, 1-4 (2023)

  24. [32]

    Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying

    Soroush, A., et al. Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying. (2024)

  25. [33]

    Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study

    Travis, Z., et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet. Digital health 6(2024). 15

  26. [34]

    & Raymond H, M

    Danielle S, B., Hugo JWL, A. & Raymond H, M. Approaching autonomy in medical artificial intelligence. The Lancet. Digital health 2(2020)

  27. [35]

    Tackling bias in AI health datasets through the STANDING Together initiative

    Shaswath, G., et al. Tackling bias in AI health datasets through the STANDING Together initiative. Nature medicine 28(2022)

  28. [36]

    & Joseph C, K

    Mirja, M., Marium M, R. & Joseph C, K. Bias in AI-based models for medical applications: challenges and mitigation strategies. NPJ digital medicine 6(2023)

  29. [37]

    Digital Health to Patient-Facing Artificial Intelligence: Ethical Implications and Threats to Dignity for Patients With Cancer

    Amar H, K., et al. Digital Health to Patient-Facing Artificial Intelligence: Ethical Implications and Threats to Dignity for Patients With Cancer. JCO oncology practice 20(2024)

  30. [38]

    Food and Drug Administration. U.S.A. Focus Areas of Regulatory Science - Introduction | FDA. Accessed through: https://www.fda.gov/science-research/focus-areas-regulatory-science- report/focus-areas-regulatory-science-introduction (2024)

  31. [39]

    & Farid, S.S

    Grandis, G.D., Brass, I. & Farid, S.S. Is regulatory innovation fit for purpose? A case study of adaptive regulation for advanced biotherapeutics. Regulation & Governance (2023)

  32. [40]

    From adaptive licensing to adaptive pathways: delivering a flexible life-span approach to bring new drugs to patients

    H-G, E., et al. From adaptive licensing to adaptive pathways: delivering a flexible life-span approach to bring new drugs to patients. Clinical pharmacology and therapeutics 97(2015)

  33. [41]

    & Pall, J

    Emily, L., Dalia, D., Jacoline, B. & Pall, J. The Sandbox Approach and its Potential for Use in Health Technology Assessment: A Literature Review. Applied health economics and health policy 19(2021)

  34. [42]

    Privacy Enhancing Technology Sandboxes - Infocomm Media Development Authority

    Infocomm Media Development Authority, Singapore. Privacy Enhancing Technology Sandboxes - Infocomm Media Development Authority. Accessed through: https://www.imda.gov.sg/how-we- can-help/data-innovation/privacy-enhancing-technology-sandboxes (2024)

  35. [43]

    Regulatory sandboxes in artificial intelligence

    OECD. Regulatory sandboxes in artificial intelligence. Vol. No. 356 (OECD Digital Economy Papers, 2023)

  36. [44]

    LLM-based agentic systems in medicine and healthcare

    Qiu, J., et al. LLM-based agentic systems in medicine and healthcare. Nature Machine Intelligence 6, 1418-1420 (2024)

  37. [45]

    Evaluating large language models as agents in the clinic

    Mehandru, N., et al. Evaluating large language models as agents in the clinic. npj Digital Medicine 7, 1-3 (2024)

  38. [46]

    (@OpenAI, 2024)

    Introducing OpenAI o1. (@OpenAI, 2024)

  39. [47]

    Learning to Reason with LLMs

    OpenAI. Learning to Reason with LLMs. Vol. 16 December 2024 (https://openai.com/index/learning-to-reason-with-llms/, September 12, 2024)

  40. [48]

    Regulating the Future of Health: CoRE’s Tenth Anniversary Perspective

    John CW Lim, Tan Koi Wei Chuen & Vogel, S. Regulating the Future of Health: CoRE’s Tenth Anniversary Perspective. in CoRE Regulatory Perspective, Vol. 2024 (Duke-NUS Medical School, Centre of Regulatory Excellence, https://www.duke-nus.edu.sg/core/think-tank/core-regulatory- p...

  41. [49]

    CrowdStrike Technology Outage Causing Global Disruption, Including Impacts on Hospitals | AHA

    Association, A.H. CrowdStrike Technology Outage Causing Global Disruption, Including Impacts on Hospitals | AHA. (@ahahospitals, https://www.aha.org/advisory/2024-07-19-crowdstrike- technology-outage-causing-global-disruption-across-industries-including-impacts-hospitals-and, ...

  42. [50]

    A Review of Embedded Machine Learning Based on Hardware, Application, and Sensing Scheme

    Biglari A & Wei, T. A Review of Embedded Machine Learning Based on Hardware, Application, and Sensing Scheme. Sensors (Basel) 4, 2131 (2023)

  43. [51]

    & Gilbert, S

    Riedemann, L., Labonne, M. & Gilbert, S. The path forward for large language models in medicine is open. npj Digit. Med. 7, 339 (2024)

  44. [52]

    & Sang-Soo, L

    Chiranjib, C., Manojit, B. & Sang-Soo, L. Artificial intelligence enabled ChatGPT and large language models in drug target discovery, drug discovery, and development. Molecular therapy. Nucleic acids 33(2023)

  45. [53]

    Use of Artificial Intelligence in Drug Development

    Druedahl, L.C., et al. Use of Artificial Intelligence in Drug Development. JAMA Network Open 7(2024)

  46. [54]

    & Antonio, L

    Amit, G. & Antonio, L. Unleashing the power of generative AI in drug discovery. Drug discovery today 29(2024)

  47. [55]

    ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries

    Swanson, K., et al. ADMET-AI: a machine learning ADMET platform for evaluation of large-scale chemical libraries. Bioinformatics 40(2024)

  48. [56]

    Generative AI and Large Language Models in Reducing Medication Related Harm and Adverse Drug Events – A Scoping Review

    Ong, J.C.L., et al. Generative AI and Large Language Models in Reducing Medication Related Harm and Adverse Drug Events – A Scoping Review. medRxiv. Accessed: https://doi.org/10.1101/2024.09.13.24313606 (2024)

  49. [57]

    Working Groups: Artificial Intelligence/Machine Learning-enabled

    International Medical Device Regulators Forum. Working Groups: Artificial Intelligence/Machine Learning-enabled. Vol. 2024 (https://www.imdrf.org/working-groups/artificial-intelligencemachine- learning-enabled)

  50. [58]

    A framework for human evaluation of large language models in healthcare derived from literature review

    Thomas YC, T., et al. A framework for human evaluation of large language models in healthcare derived from literature review. NPJ digital medicine 7(2024). 16

  51. [59]

    Assessment of Adherence to Reporting Guidelines by Commonly Used Clinical Prediction Models From a Single Vendor: A Systematic Review

    Jonathan H, L., et al. Assessment of Adherence to Reporting Guidelines by Commonly Used Clinical Prediction Models From a Single Vendor: A Systematic Review. JAMA network open 5(2022)

  52. [60]

    The value of standards for health datasets in artificial intelligence-based applications

    Anmol, A., et al. The value of standards for health datasets in artificial intelligence-based applications. Nature medicine 29(2023)

  53. [61]

    G Gallifant, J., Afshar, M., Ameen, S. et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med (2025)

  54. [62]

    Cruz Rivera, S., Liu, X., Chan, AW. et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med 26, 1351–1363 (2020)

  55. [63]

    Liu, X., Cruz Rivera, S., Moher, D. et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med 26, 1364–1374 (2020)

  56. [64]

    Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist

    Yilin, N., et al. Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist. The Lancet. Digital health (2024)

  57. [65]

    An ethics assessment tool for artificial intelligence implementation in healthcare: CARE-AI

    Yilin, N., et al. An ethics assessment tool for artificial intelligence implementation in healthcare: CARE-AI. Nature medicine (2024)

  58. [66]

    A Nationwide Network of Health AI Assurance Laboratories

    Nigam H, S., et al. A Nationwide Network of Health AI Assurance Laboratories. JAMA 331(2024)

  59. [67]

    UC Davis Health: Health and Technology

    Swanson, T. UC Davis Health: Health and Technology. UC Davis Health, NODE.health, and Leading Health Systems launch VALID AI. Accessed through: https://health.ucdavis.edu/news/headlines/uc-davis-health-and-leading-health-systems-launch- valid-ai/2023/10, (2023)

  60. [68]

    & Singh, K

    Price, W.N., Sendak, M., Balu, S. & Singh, K. Enabling collaborative governance of medical AI. Nature Machine Intelligence 5, 821-823 (2023)

  61. [69]

    Disparities in clinical studies of AI enabled applications from a global perspective

    Rui, Y., et al. Disparities in clinical studies of AI enabled applications from a global perspective. NPJ digital medicine 7(2024)

  62. [70]

    Considerations for addressing bias in artificial intelligence for health equity

    Michael D, A., et al. Considerations for addressing bias in artificial intelligence for health equity. NPJ digital medicine 6(2023)

  63. [71]

    & Gilbert, S

    Brückner, S., Brightwell, C. & Gilbert, S. FDA launches health care at home initiative to drive equity in digital medical care. npj Digital Medicine 7, 1-3 (2024)

  64. [72]

    Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias

    Chen, S., et al. Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias. ArXiv. Accessed through: https://doi.org/10.48550/arXiv.2405.05506 (2024)

  65. [73]

    Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study

    Zack, T., et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet. Digital health 6(2024)

  66. [74]

    The translational gap for gene therapies in low- and middle-income countries

    Kevin W, D., et al. The translational gap for gene therapies in low- and middle-income countries. Science translational medicine 16(2024)

  67. [75]

    & Sandra, B

    Tadeusz, C.-H., Ritvij, S., Miriam, A., Stephan, B. & Sandra, B. Artificial intelligence for strengthening healthcare systems in low- and middle-income countries: a systematic scoping review. NPJ digital medicine 5(2022)

  68. [76]

    Impact of Artificial Intelligence Assessment of Diabetic Retinopathy on Referral Service Uptake in a Low-Resource Setting: The RAIDERS Randomized Trial

    Wanjiku, M., et al. Impact of Artificial Intelligence Assessment of Diabetic Retinopathy on Referral Service Uptake in a Low-Resource Setting: The RAIDERS Randomized Trial. Ophthalmology science 2(2022)

  69. [77]

    Feasibility and acceptance of artificial intelligence-based diabetic retinopathy screening in Rwanda

    Noelle, W., et al. Feasibility and acceptance of artificial intelligence-based diabetic retinopathy screening in Rwanda. The British journal of ophthalmology 108(2024)

  70. [78]

    Integrated image-based deep learning and language models for primary diabetes care

    Li, J., et al. Integrated image-based deep learning and language models for primary diabetes care. Nature Medicine, 1-11 (2024)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.