Pith. sign in

REVIEW 5 major objections 6 minor 49 references

Advice for Diabetes Self-Management by ChatGPT Models: Challenges and Recommendations

T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper reports that ChatGPT 3.5, 4, 4o, and 4o mini still produce diabetes self-management advice that is often generic, assumption-laden, and potentially dangerous, and that most critiques from a 2023 study still hold.

desk verdict A useful reminder that newer ChatGPT models still fumble diabetes advice, but the 'often' claim outruns the evidence. read the letter →

arxiv 2501.07931 v1 pith:3FURCADR submitted 2025-01-14 cs.AI

classification cs.AI
keywords ChatGPTdiabetesself-managementlargelanguagemodelspatientsafetymedicaladviceevaluationretrievalaugmentedgenerationDSMEShealthcareAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that current ChatGPT models, despite strong benchmark performance, remain unsafe as standalone advisors for diabetes self-management. The authors re-run a 2023 evaluation of 20 diabetes questions on ChatGPT 3.5, 4, 4o, and 4o mini and find that most earlier critiques still hold: advice stays generic, insulin regimens are conflated, blood-glucose units are assumed without clarification, and pseudo-hypoglycemia is misread. Because some failures, such as unit mix-ups, can be life-threatening, the paper argues these models need human oversight and should be wrapped in a commonsense evaluation layer and in retrieval-augmented generation with disease-specific external memory. The care case is that millions of people managing diabetes daily might act on AI advice, so knowing whether it has improved matters.

What carries the argument

The evaluative machinery is a replication protocol: the same 20 unstructured diabetes patient questions from the 2023 baseline study, covering four DSMES domains, are posed to each ChatGPT version in a conversational style without prompt engineering, and responses are rated by two healthcare professionals, a GP and a dietician, on consistency, reliability, and accuracy, then checked against the earlier study's critiques. The proposed fixes are a commonsense evaluation layer, which would force models to run systematic checks and ask clarifying questions before responding, and an Advanced Retrieval Augmented Generation system, which would ground advice in disease-specific external memory such as current guidelines from authoritative health sources. Together these components are meant to reduce assumption-driven errors and make AI advice safer for chronic disease self-management.

What would settle it

Run the same 20 queries against current ChatGPT versions with a pre-registered rubric and blinded clinician raters; the central claim would be falsified if most 2023 critiques no longer hold, for instance if the models consistently ask whether a reading is in mg/dL or mmol/L before advising on a blood sugar of 25.

Watch

Extended reading notes

Core claim

Across 20 unstructured diabetes patient queries spanning diet, exercise, hypo- and hyperglycemia education, insulin storage, and administration, ChatGPT 3.5 and ChatGPT 4 show only slight improvement over the 2023 baseline, and most of the earlier critiques still apply to ChatGPT 4, 4o, and 4o mini. The models consistently give generalized rather than personalized advice, fail to distinguish different insulin regimens, assume blood-glucose readings are in mg/dL without asking, misclassify pseudo-hypoglycemia as hypoglycemia unawareness, and omit type-specific insulin storage guidance. One concrete example is the question 'My blood sugar is 25, what should I do?', where different ChatGPT versions disagree on whether the value is critically low or critically high because they do not verify the measurement unit. The paper concludes that these models are not safe as standalone diabetes self-management advisors and that their practical effectiveness depends on human oversight and targeted technical safeguards.

Load-bearing premise

The study assumes that the subjective ratings of two healthcare professionals on 20 pre-selected questions are a reliable and generalizable measure of ChatGPT's diabetes advice quality, with no formal scoring rubric or inter-rater reliability reported.

Editorial extensions

If this is right

  • If the paper is right, ChatGPT should not be used as a standalone diabetes self-management tool; any deployment needs a human clinician in the loop, especially for emergency and medication questions.
  • Persistent failure modes across model versions imply that scaling up training data and parameters alone will not close the safety gap; targeted interventions like clarification prompts and external knowledge retrieval are needed.
  • Patients and diabetes educators should treat model outputs as general information rather than personalized care plans, particularly for insulin dosing, snacking advice, and insulin storage.
  • A risk-tiered interaction framework, in which high-risk AI advice is validated by clinicians before reaching patients, becomes a plausible baseline for integrating healthcare LLMs safely.
  • Re-running the same 20-question critique set on each new ChatGPT version would provide a concrete benchmark for tracking whether diabetes advice safety actually improves.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The unit-assumption failure is probably one instance of a general safety gap: the same 'assume, don't ask' pattern may appear whenever units, drug names, or treatment conventions vary by region, so similar dangers could exist for other chronic conditions.
  • A simple testable remedy implied by the findings is to require models to ask a clarifying question before any numeric or medication advice; a 'clarification rate' measured on this 20-question set could serve as a safety benchmark.
  • The cultural and economic analysis suggests that even an Advanced RAG system will not fix equity unless the external knowledge base itself contains region-specific dietary and cost data; otherwise retrieval may just anchor Western defaults.
  • One could test the authors' proposed fix directly by running the same 20 questions through a RAG-augmented ChatGPT and having clinicians rate it with the same rubric; if most critiques vanish, the recommendation would be validated.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper evaluates diabetes self-management advice generated by ChatGPT models (primarily GPT-3.5 and GPT-4, with additional claims about GPT-4o and GPT-4o mini) by posing 20 unstructured queries covering diet, exercise, hypoglycemia/hyperglycemia, and insulin storage and administration. Responses were reviewed by two healthcare professionals using qualitative criteria of consistency, reliability, and accuracy. The authors report that the latest models show only slight improvement over GPT-3.5, that most critiques from a 2023 study by Sng et al. still apply, and that the models often give advice without seeking clarification, which can be dangerous (e.g., misinterpreting 'blood sugar is 25' as mg/dL rather than mmol/L). The paper then proposes a commonsense evaluation layer and an advanced Retrieval Augmented Generation (RAG) framework as mitigations, without implementing or testing them.

Significance. If the empirical claims were robust, the paper would be a useful, timely warning that ChatGPT models remain unsafe as standalone diabetes self-management advisors. The authors build on prior work (Sng et al. 2023) and provide concrete, illustrative examples of dangerous misinterpretations, which is valuable for the AI-in-healthcare community. The paper also draws attention to important issues such as unit-of-measurement ambiguity, cultural insensitivity in meal planning, and non-English support. However, the evidentiary basis is thin: single unrepeated runs, no scoring rubric, no inter-rater reliability, and no released transcripts. The central frequency claim ('often') is therefore not statistically supported, and the proposed solutions are untested. The significance is conditional on a more rigorous evaluation.

major comments (5)
  1. [Section 3 and Section 4.3.3, Table 5] The central claim that 'both models often provide advice without seeking necessary clarification' is not supported by the reported methodology. Each query was run once per model with no reported sampling temperature, no repeated sampling, and no released transcripts. Since LLM outputs are stochastic, a single draw cannot establish a frequency such as 'often' without a variance estimate or confidence interval. The authors should either provide repeated runs with full response logs, or weaken the claim to 'in the instances we observed.' The paper's own Section 5.3 concedes that output varies with prompt modifications and model updates, which makes this methodological gap load-bearing.
  2. [Section 3, 'We evaluated the responses...'] The evaluation by two healthcare professionals is described without a scoring rubric, without operational definitions for 'consistency,' 'reliability,' and 'accuracy,' and without any measure of inter-rater agreement (e.g., Cohen's kappa). Tables 1 and 2 present qualitative summaries for only two questions, and the overall claim of 'slight improvement' is not quantified. Without a rubric and agreement measure, the ratings cannot be distinguished from subjective impressions, and the cross-version comparison is not reproducible.
  3. [Section 4.3.3, Table 5 vs. Section 3] There is a mismatch between the models described in the methodology and those in the results. Section 3 states that the study posed 20 queries to ChatGPT 3.5 and ChatGPT-4 (text only), yet Table 5 reports responses from ChatGPT 4o and ChatGPT 4o mini. The abstract also refers only to 'ChatGPT versions 3.5 and 4,' while the introduction claims evaluation of GPT-4o and GPT-4o mini. The authors must state which models were tested, when, and under what interface/version, and ensure the abstract, introduction, methodology, and results are consistent.
  4. [Tables 3 and 4] Tables 3 and 4 appear to be near-duplicates but contain inconsistent ratings. For the insulin pen priming row, Table 3 rates the critique as 'Partially Fair' with severity 'Moderate,' while Table 4 rates it 'Unfair' with severity 'None.' The severity legends also differ (Table 3 includes 'Low,' Table 4 includes 'None'). This inconsistency undermines the reliability of the critique evaluation and suggests an editing error. The authors should consolidate the tables and verify that all ratings are consistent.
  5. [Section 6, 'Commonsense Evaluation Layer' and 'Advice Improvements with RAG'] The proposed commonsense evaluation layer and the advanced RAG-based chronic disease management model are presented as key contributions, but they are not implemented, evaluated, or validated anywhere in the paper. There is no architecture detail beyond a generic figure, no data, and no experiment. As a recommendation or future-work section this is acceptable, but as stated in the contributions it overstates the paper's deliverable. The authors should clearly label these as untested proposals and either remove them from the contribution list or provide a proof-of-concept evaluation.
minor comments (6)
  1. [Section 3, 'Building on the work of Sng et al.'] The paper says the 20 questions were 'noted in [44]' but does not reproduce the full question set. Since the questions are central to reproducibility, they should be included in an appendix or supplementary material.
  2. [Section 5.2.1, Table 6 and Table 7] The cost estimates in Table 6 and Table 7 mix Australian dollars and Pakistani rupees without a stated exchange-rate date or source. Please add a conversion reference or a note that prices are approximate.
  3. [Section 4.1, Tables 1 and 2] The quoted examples in Tables 1 and 2 are not time-stamped or version-stamped beyond the model name. Since ChatGPT responses change over time, please provide the exact access dates and, if possible, the conversation IDs or saved transcripts.
  4. [Section 7, Limitations] The limitations section mentions simulated patient inquiries and the narrow query range, but it does not acknowledge the lack of repeated sampling, the absence of inter-rater reliability, or the absence of a scoring rubric. These are the main methodological limitations and should be discussed.
  5. [Abstract] The abstract says 'both models often provide advice without seeking necessary clarification,' but the paper does not quantify how many of the 20 responses exhibited this behavior. Please report a count or percentage, or revise the wording.
  6. [Section 1, Figure 1] Figure 1 is referenced as showing 'good and bad advice from ChatGPT 4,' but the figure itself is not included in the text and the caption is missing. Either include the figure or remove the reference.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper reports an empirical evaluation with an inherited but external methodology, and its recommendations are untested proposals rather than fitted predictions.

full rationale

The paper is an empirical evaluation study rather than a derivation, and no claimed result reduces to its own inputs by construction. The methodology is explicitly adapted from Sng et al. [44], but that prior work is external to the present authors, so the reused critique framework is inherited evidence, not a self-citation chain. The central claims about ChatGPT's diabetes advice quality are based on 20 queries rated by two healthcare professionals; these ratings are not fitted parameters, and the reported failures are observed outputs, not quantities derived from the evaluation criteria. The abstract's statement that the models 'often provide advice without seeking necessary clarification' is an empirical frequency claim whose reliability may be questioned because of single unrepeated runs and subjective ratings, but that is a methodological limitation about generalizability, not circularity. The proposed commonsense evaluation layer and Advanced RAG recommendations are explicitly future directions, not validated outputs of the study, so they cannot be circular predictions. No equation, definition, or parameter in the paper is constructed so that a later claimed result is identical to an earlier input. Accordingly, the paper is self-contained in the sense relevant to circularity analysis, and no circular step is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 2 invented entities

The paper rests on three domain assumptions: the representativeness of the 20 questions, the validity of two subjective expert raters, and the stability/representativeness of sampled ChatGPT responses. It also proposes two untested system components (commonsense evaluation layer and Advanced RAG). There are no fitted parameters.

assumptions (3)
  • domain assumption The 20 DSMES questions from Sng et al. (2023) are representative of the range of diabetes self-management queries.
    The study uses these questions without justification that they cover the breadth of patient needs; Section 3.
  • domain assumption Subjective ratings by one general practitioner and one dietician are a valid and sufficient measure of advice quality.
    No inter-rater reliability, no scoring rubric, and only two raters; Section 3.
  • domain assumption The ChatGPT responses obtained at the time of testing are stable and representative of each model version's behavior.
    LLMs are non-deterministic and can vary by date and prompt; no repeated sampling was reported.
invented entities (2)
  • Commonsense evaluation layer
    purpose: A proposed add-on to system cards to check and clarify AI responses before giving medical advice.
    Described in Section 6.1 but not implemented or evaluated in this study.
  • Advanced RAG-based chronic disease management model
    purpose: Proposed architecture combining an LLM with external medical databases to personalize and update advice.
    Described in Section 6.2 with a figure but no prototype, test, or benchmark.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advice for Diabetes Self-Management by ChatGPT Models: Challenges and Recommendations." pith.science (2026). https://pith.science/paper/3FURCADR

@misc{pith2026250107931,
  author       = {Pith},
  title        = {Pith review of: Advice for Diabetes Self-Management by ChatGPT Models: Challenges and Recommendations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3FURCADR}},
  note         = {Machine review of arXiv:2501.07931}
}
read the original abstract

Given their ability for advanced reasoning, extensive contextual understanding, and robust question-answering abilities, large language models have become prominent in healthcare management research. Despite adeptly handling a broad spectrum of healthcare inquiries, these models face significant challenges in delivering accurate and practical advice for chronic conditions such as diabetes. We evaluate the responses of ChatGPT versions 3.5 and 4 to diabetes patient queries, assessing their depth of medical knowledge and their capacity to deliver personalized, context-specific advice for diabetes self-management. Our findings reveal discrepancies in accuracy and embedded biases, emphasizing the models' limitations in providing tailored advice unless activated by sophisticated prompting techniques. Additionally, we observe that both models often provide advice without seeking necessary clarification, a practice that can result in potentially dangerous advice. This underscores the limited practical effectiveness of these models without human oversight in clinical settings. To address these issues, we propose a commonsense evaluation layer for prompt evaluation and incorporating disease-specific external memory using an advanced Retrieval Augmented Generation technique. This approach aims to improve information quality and reduce misinformation risks, contributing to more reliable AI applications in healthcare settings. Our findings seek to influence the future direction of AI in healthcare, enhancing both the scope and quality of its integration.

Figures

Figures reproduced from arXiv: 2501.07931 by the authors.

Figure 1
Figure 1. Good and bad advice from ChatGPT 4 The advancements and proliferation of Large Lan￾guage Models (LLMs) such as ChatGPT, with advanced conversational abilities, extensive medical knowledge, and proficiency in scenario-based learning, mark them as promising tools for patient advice in health manage￾ment. ChatGPT-4, in particular, has shown strong per￾formance on medical benchmarks like the United States Medical Licens… view at source ↗
Figure 2
Figure 2. Enhanced Adaptability and integration of [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 44 canonical work pages

  1. [1]

    Knowledge-Infused LLM-Powered Conversational Health Agent: A Case Study for Diabetes Patients

    Mahyar Abbasian, Zhongqi Yang, Elahe Khatibi, Pengfei Zhang, Nitish Nagesh, Iman Azimi, Ramesh Jain, and Amir M Rahmani. Knowledge-infused llm- powered conversational health agent: A case study for diabetes patients. arXiv preprint arXiv:2402.10153 , 2024

  2. [2]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal 11 Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  3. [3]

    Diabetes in australia, 2024

    Diabetes Australia. Diabetes in australia, 2024. Ac- cessed: February 6, 2024

  4. [4]

    Comparing physician and artificial intel- ligence chatbot responses to patient questions posted to a public social media forum

    John W Ayers, Adam Poliak, Mark Dredze, Eric C Leas, Zechariah Zhu, Jessica B Kelley, Dennis J Faix, Aaron M Goodman, Christopher A Longhurst, Michael Hogarth, et al. Comparing physician and artificial intel- ligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine , 2023

  5. [5]

    Credibility of chatgpt in the assessment of obesity in type 2 diabetes accord- ing to the guidelines

    Tugba Barlas, Alev Eroglu Altinova, Mujde Akturk, and Fusun Balos Toruner. Credibility of chatgpt in the assessment of obesity in type 2 diabetes accord- ing to the guidelines. International Journal of Obesity , 48(2):271–275, 2024

  6. [6]

    On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021

    Emily M Bender, Timnit Gebru, Angelina McMillan- Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021

  7. [7]

    A meta-analysis of the effec- tiveness of health belief model variables in predicting behavior

    Christopher J Carpenter. A meta-analysis of the effec- tiveness of health belief model variables in predicting behavior. Health communication, 25(8):661–669, 2010

  8. [8]

    Evaluation of gpt-3.5 and gpt-4 for supporting real-world information needs in health- care delivery

    Debadutta Dash, Rahul Thapa, Juan M Banda, Ak- shay Swaminathan, Morgan Cheatham, Mehr Kashyap, Nikesh Kotecha, Jonathan H Chen, Saurabh Gombar, Lance Downing, et al. Evaluation of gpt-3.5 and gpt-4 for supporting real-world information needs in health- care delivery. arXiv preprint arXiv:2304.13714 , 2023

Show all 49 references
  1. [9]

    Exploring capabilities of large language mod- els such as chatgpt in radiation oncology

    Fabio Dennst¨ adt, Janna Hastings, Paul Martin Putora, Erwin Vu, Galina F Fischer, Krisztian S¨ uveg, Markus Glatzer, Elena Riggenbach, Hˆ ong-Linh H` a, and Nikola Cihoric. Exploring capabilities of large language mod- els such as chatgpt in radiation oncology. Advances in ra...

  2. [10]

    Amit Kumar Dey. Chatgpt in diabetes care: An overview of the evolution and potential of generative artificial intelligence model like chatgpt in augmenting clinical and patient outcomes in the management of diabetes. International Journal of Diabetes and Tech- nology, 2(2):66–72, 2023

  3. [11]

    Exploring the clinical translation of generative models like chatgpt: promise and pitfalls in radiology, from patients to population health

    Florence X Doo, Tessa S Cook, Eliot L Siegel, Anupam Joshi, Vishwa Parekh, Ameena Elahi, and H Yi Paul. Exploring the clinical translation of generative models like chatgpt: promise and pitfalls in radiology, from patients to population health. Journal of the American College ...

  4. [12]

    The role of chatgpt, gen- erative language models, and artificial intelligence in medical education: a conversation with chatgpt and a call for papers

    Gunther Eysenbach et al. The role of chatgpt, gen- erative language models, and artificial intelligence in medical education: a conversation with chatgpt and a call for papers. JMIR Medical Education , 9(1):e46885, 2023

  5. [13]

    Enhancing retrieval processes for language gener- ation with augmented queries

    Julien Pierre Edmond Ghali, Kosuke Shima, Koichi Moriyama, Atsuko Mutoh, and Nobuhiro Inuzuka. Enhancing retrieval processes for language gener- ation with augmented queries. arXiv preprint arXiv:2402.16874, 2024

  6. [14]

    Gilson, C

    A. Gilson, C. Safranek, T. Huang, V. Socrates, L. Chi, R. A. Taylor, and D. Chartash. How well does chatgpt do when taking the medical licensing exams? the impli- cations of large language models for medical education and knowledge assessment. medRxiv, 2022. 2022-12

  7. [15]

    May chatgpt be a tool producing medical information for common inflamma- tory bowel disease patients’ questions? an evidence- controlled analysis

    Antonietta Gerarda Gravina, Raffaele Pellegrino, Ma- rina Cipullo, Giovanna Palladino, Giuseppe Imperio, Andrea Ventura, Salvatore Auletta, Paola Ciamarra, and Alessandro Federico. May chatgpt be a tool producing medical information for common inflamma- tory bowel disease pati...

  8. [16]

    A new taxonomy for technology-enabled diabetes self-management interven- tions: results of an umbrella review

    Deborah A Greenwood, Michelle L Litchman, Diana Isaacs, Julia E Blanchette, Jane K Dickinson, Allyson Hughes, Vanessa D Colicchio, Jiancheng Ye, Kirsten Yehl, Andrew Todd, et al. A new taxonomy for technology-enabled diabetes self-management interven- tions: results of an umbr...

  9. [17]

    National stan- dards for diabetes self-management education and sup- port

    Linda Haas, Melinda Maryniuk, Joni Beck, Carla E Cox, Paulina Duker, Laura Edwards, Ed Fisher, Lenita Hanson, Daniel Kent, Leslie Kolb, et al. National stan- dards for diabetes self-management education and sup- port. The Diabetes Educator , 38(5):619–629, 2012

  10. [18]

    Artificial intelligence for clinical trial de- sign

    Stefan Harrer, Pratik Shah, Bhavna Antony, and Jiany- ing Hu. Artificial intelligence for clinical trial de- sign. Trends in pharmacological sciences , 40(8):577– 591, 2019

  11. [19]

    Do we still need clinical language models? In Con- ference on Health, Inference, and Learning , pages 578–

    Evan Hernandez, Diwakar Mahajan, Jonas Wulff, Micah J Smith, Zachary Ziegler, Daniel Nadler, Pe- ter Szolovits, Alistair Johnson, Emily Alsentzer, et al. Do we still need clinical language models? In Con- ference on Health, Inference, and Learning , pages 578–

  12. [20]

    A generative pretrained transformer (gpt)–powered chat- bot as a simulated patient to practice history taking: Prospective, mixed methods study

    Friederike Holderried, Christian Stegemann-Philipps, Lea Herschbach, Julia-Astrid Moldt, Andrew Nevins, Jan Griewatz, Martin Holderried, Anne Herrmann- Werner, Teresa Festl-Wietek, Moritz Mahling, et al. A generative pretrained transformer (gpt)–powered chat- bot as a simulate...

  13. [21]

    Chatgpt and antimicrobial advice: the end of the con- sulting infection doctor? The Lancet Infectious Dis- eases, 23(4):405–406, 2023

    Alex Howard, William Hope, and Alessandro Gerada. Chatgpt and antimicrobial advice: the end of the con- sulting infection doctor? The Lancet Infectious Dis- eases, 23(4):405–406, 2023

  14. [22]

    Chatgpt- versus human-generated answers to frequently asked questions about diabetes: A turing test-inspired sur- vey among employees of a danish diabetes center

    Adam Hulman, Ole Lindg˚ ard Dollerup, Jesper Friis Mortensen, Matthew E Fenech, Kasper Norman, Hen- rik Støvring, and Troels Krarup Hansen. Chatgpt- versus human-generated answers to frequently asked questions about diabetes: A turing test-inspired sur- vey among employees of ...

  15. [23]

    Twelve tips to lever- age ai for efficient and effective medical question gen- eration: A guide for educators using chat gpt

    Inthrani Raja Indran, Priya Paramanathan, Neelima Gupta, and Nurulhuda Mustafa. Twelve tips to lever- age ai for efficient and effective medical question gen- eration: A guide for educators using chat gpt. Medical Teacher, pages 1–6, 2023

  16. [24]

    The accuracy and potential racial and ethnic biases of gpt-4 in the diagnosis and triage of health conditions: Evaluation study

    Naoki Ito, Sakina Kadomatsu, Mineto Fujisawa, Kiyomitsu Fukaguchi, Ryo Ishizawa, Naoki Kanda, Daisuke Kasugai, Mikio Nakajima, Tadahiro Goto, and Yusuke Tsugawa. The accuracy and potential racial and ethnic biases of gpt-4 in the diagnosis and triage of health conditions: Eval...

  17. [25]

    Chatgpt makes medicine easy to swallow: an exploratory case study on simplified ra- diology reports

    Katharina Jeblick, Balthasar Schachtner, Jakob Dexl, Andreas Mittermeier, Anna Theresa St¨ uber, Johanna 12 Topalis, Tobias Weber, Philipp Wesp, Bastian Oliver Sabel, Jens Ricke, et al. Chatgpt makes medicine easy to swallow: an exploratory case study on simplified ra- diology...

  18. [26]

    Assessing the accuracy and reliability of ai-generated medical responses: an evaluation of the chat-gpt model

    Douglas Johnson, Rachel Goodman, J Patrinely, Cosby Stone, Eli Zimmerman, Rebecca Donald, Sam Chang, Sean Berkowitz, Avni Finn, Eiman Jahangir, et al. Assessing the accuracy and reliability of ai-generated medical responses: an evaluation of the chat-gpt model. Research square, 2023

  19. [27]

    Using chatgpt to predict the future of diabetes technology

    David Kerr and David C Klonoff. Using chatgpt to predict the future of diabetes technology. Journal of Diabetes Science and Technology , 1:2, 2023

  20. [28]

    Can chatgpt help in the awareness of diabetes? Annals of Biomedical Engineering, 51(10):2125–2129, 2023

    Imran Khan and Rashi Agarwal. Can chatgpt help in the awareness of diabetes? Annals of Biomedical Engineering, 51(10):2125–2129, 2023

  21. [29]

    Performance of chatgpt on usmle: potential for ai-assisted medical education using large language models

    Tiffany H Kung, Morgan Cheatham, Arielle Mede- nilla, Czarina Sillos, Lorie De Leon, Camille Elepa˜ no, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al. Performance of chatgpt on usmle: potential for ai-assisted medical education using large language m...

  22. [30]

    Trustworthy artificial intelligence and the european union ai act: On the conflation of trustworthiness and acceptability of risk

    Johann Laux, Sandra Wachter, and Brent Mittelstadt. Trustworthy artificial intelligence and the european union ai act: On the conflation of trustworthiness and acceptability of risk. Regulation & Governance , 18(1):3–32, 2024

  23. [31]

    J. Li, J. Huang, L. Zheng, and X. Li. Application of artificial intelligence in diabetes education and manage- ment: present status and promising prospect. Frontiers in Public Health , 8:173, 2020

  24. [32]

    Cost-effectiveness of interventions to prevent and control diabetes mellitus: a systematic review

    Rui Li, Ping Zhang, Lawrence E Barker, Farah M Chowdhury, and Xuanping Zhang. Cost-effectiveness of interventions to prevent and control diabetes mellitus: a systematic review. Diabetes care, 33(8):1872–1894, 2010

  25. [33]

    Can large language models reason about medical questions? Pat- terns, 2023

    Valentin Li´ evin, Christoffer Egeberg Hother, An- dreas Geert Motzfeldt, and Ole Winther. Can large language models reason about medical questions? Pat- terns, 2023

  26. [34]

    Benchmarking large language models’ performances for myopia care: a comparative analysis of chatgpt-3.5, chatgpt-4.0, and google bard

    Zhi Wei Lim, Krithi Pushpanathan, Samantha Min Er Yew, Yien Lai, Chen-Hsin Sun, Janice Sing Harn Lam, David Ziyou Chen, Jocelyn Hui Lin Goh, Mar- cus Chun Jin Tan, Bin Sheng, et al. Benchmarking large language models’ performances for myopia care: a comparative analysis of cha...

  27. [35]

    Translating radiology re- ports into plain language using chatgpt and gpt-4 with prompt learning: results, limitations, and potential

    Qing Lyu, Josh Tan, Michael E Zapadka, Janardhana Ponnatapura, Chuang Niu, Kyle J Myers, Ge Wang, and Christopher T Whitlow. Translating radiology re- ports into plain language using chatgpt and gpt-4 with prompt learning: results, limitations, and potential. Visual Computing ...

  28. [36]

    Empowering personalized pharmacogenomics with gen- erative ai solutions

    Mullai Murugan, Bo Yuan, Eric Venner, Christie M Ballantyne, Katherine M Robinson, James C Coons, Liwen Wang, Philip E Empey, and Richard A Gibbs. Empowering personalized pharmacogenomics with gen- erative ai solutions. medRxiv, pages 2024–02, 2024

  29. [37]

    Capabilities of gpt-4 on medical challenge problems

    Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375, 2023

  30. [38]

    Gpt-4o system card

    OpenAI. Gpt-4o system card. OpenAI website, 2024. Accessed: 2024-08-10

  31. [39]

    Is chatgpt an effective tool for providing dietary advice? Nutrients, 16(4):469, 2024

    Valentina Ponzo, Ilaria Goitre, Enrica Favaro, Fabio Dario Merlo, Maria Vittoria Mancino, Sergio Riso, and Simona Bo. Is chatgpt an effective tool for providing dietary advice? Nutrients, 16(4):469, 2024

  32. [40]

    Margaret A Powers, Joan Bardsley, Marjorie Cypress, Paulina Duker, Martha M Funnell, Amy Hess Fischl, Melinda D Maryniuk, Linda Siminerio, and Eva Vi- vian. Diabetes self-management education and sup- port in type 2 diabetes: a joint position statement of the american diabetes...

  33. [41]

    Retrieval augmented chest x-ray report generation using openai gpt models

    Mercy Ranjit, Gopinath Ganapathy, Ranjit Manuel, and Tanuja Ganu. Retrieval augmented chest x-ray report generation using openai gpt models. In Ma- chine Learning for Healthcare Conference , pages 650–

  34. [42]

    A critical review of chat- gpt as a potential substitute for diabetes educators

    Samriddhi Sharma, Sandhya Pajai, Roshan Prasad, Mayur B Wanjari, Pratiksha K Munjewar, Ranjana Sharma, and Aniket Pathade. A critical review of chat- gpt as a potential substitute for diabetes educators. Cureus, 15(5), 2023

  35. [43]

    Chatgpt and other large language models are double- edged swords

    Yiqiu Shen, Laura Heacock, Jonathan Elias, Keith D Hentel, Beatriu Reig, George Shih, and Linda Moy. Chatgpt and other large language models are double- edged swords. Radiology, 307(2):e230163, 2023

  36. [44]

    Potential and pitfalls of chatgpt and natural-language artificial intel- ligence models for diabetes education

    Gerald Gui Ren Sng, Joshua Yi Min Tung, Daniel Yan Zheng Lim, and Yong Mong Bee. Potential and pitfalls of chatgpt and natural-language artificial intel- ligence models for diabetes education. Diabetes Care, 46(5):e103–e105, 2023

  37. [45]

    Large language models in medicine

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. Large language models in medicine. Nature medicine, 29(8):1930–1940, 2023

  38. [46]

    Efficient healthcare with large language mod- els: optimizing clinical workflow and enhancing patient care

    Satvik Tripathi, Rithvik Sukumaran, and Tessa S Cook. Efficient healthcare with large language mod- els: optimizing clinical workflow and enhancing patient care. Journal of the American Medical Informatics As- sociation, page ocad258, 2024

  39. [47]

    Chatgpt: Is this version good for healthcare and re- search? Diabetes & Metabolic Syndrome: Clinical Re- search & Reviews , 17(4):102744, 2023

    Raju Vaishya, Anoop Misra, and Abhishek Vaish. Chatgpt: Is this version good for healthcare and re- search? Diabetes & Metabolic Syndrome: Clinical Re- search & Reviews , 17(4):102744, 2023

  40. [48]

    Can gpt improve the state of prior authorization via guideline based automated question answering? arXiv preprint arXiv:2402.18419, 2024

    Shubham Vatsal, Ayush Singh, and Shabnam Tafreshi. Can gpt improve the state of prior authorization via guideline based automated question answering? arXiv preprint arXiv:2402.18419, 2024

  41. [49]

    Exploring the po- tential of large language models in personalized dia- betes treatment strategies

    Hao Yang, Jiaxi Li, Siru Liu, Lei Du, Xiali Liu, Yong Huang, Qingke Shi, and Jialin Liu. Exploring the po- tential of large language models in personalized dia- betes treatment strategies. medRxiv, pages 2023–06, 2023. 13

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.