{"id":"fae8cbf5-0a87-44c6-8952-5c4483a1077b","arxiv_id":"2501.07931","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ChatGPT models still give risky, generic diabetes advice and often fail to ask clarifying questions, with only slight improvement from GPT-3.5 to GPT-4 and GPT-4o.","lead":"This study asked ChatGPT 3.5, 4, and 4o variants for diabetes self-management advice and found persistent errors, such as misreading blood sugar units and offering generic meal plans. The authors conclude the models are not safe without human oversight and propose a commonsense evaluation layer plus retrieval-augmented generation to reduce risks.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that ChatGPT models 'often' give advice without clarification rests on single, unrepeated runs; LLM stochasticity could make the observed failure rate unrepresentative.","rationale":"The reader's weakest assumption correctly identified the lack of inter-rater reliability and repeated sampling in Section 3 as a support problem. My concern is more pointed: the central claim that models 'often' fail to ask for clarification is a frequency claim, and the study design—one run per question per model—cannot estimate a frequency. This is not a minor methodological gap but a missing foundation for the headline safety conclusion. The paper's own acknowledgment of prompt sensitivity (Section 5.3) makes the single-run evidence especially fragile. The concrete test I propose would directly measure the frequency of clarification-seeking and unit-assumption behavior across repeated samples, settling whether the observed pattern is robust. I do not think this changes the reader's verdict from CONDITIONAL: the condition should simply include mandatory repeated-sampling evidence and release of transcripts. The paper's qualitative examples remain suggestive, but the 'often' quantifier and the 'potentially dangerous' generalization require this additional support.","tokens_in":17340,"tokens_out":2070,"duration_ms":22450,"concrete_test":"Run all 20 DSMES questions 30 times each (with default temperature, identical prompt wording) on ChatGPT 3.5, GPT-4, GPT-4o, and GPT-4o mini, recording for each response whether the model asks a clarifying question before providing advice, and specifically whether the 'My blood sugar is 25, what should I do?' prompt elicits a stated unit assumption. Compute the per-model proportion of clarification-seeking responses and a 95% confidence interval. If any model asks for clarification in more than 20% of runs, or if the 'blood sugar is 25' interpretation flips between mg/dL and mmol/L across runs, the paper's 'often' claim needs substantial qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central safety claim is that 'both models often provide advice without seeking necessary clarification, a practice that can result in potentially dangerous advice' (Section 4.3.3, Table 5). The evidence for this is a single set of 20 conversational queries per model (Section 3), with no repeated sampling, no reported temperature or sampling parameters, and no released transcripts. LLMs are stochastic: the same prompt can yield materially different responses across runs, and Section 5.3 itself concedes that 'output can significantly vary with slight prompt modifications or through its ongoing updates.' With one observation per question per model, the frequency word 'often' is unsupported—the observed failures could be one draw from a distribution, and the specific quoted examples (e.g., 'blood sugar is 25') may be cherry-picked or non-reproducible. The paper also reports subjective ratings from two healthcare professionals without inter-rater reliability or a scoring rubric, but the more fundamental gap is that the central frequency claim has no variance estimate at all. The 'dangerous advice' conclusion depends on 'often' being a stable property of the models, not an artifact of a single unlucky run.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates diabetes self-management advice generated by ChatGPT models (primarily GPT-3.5 and GPT-4, with additional claims about GPT-4o and GPT-4o mini) by posing 20 unstructured queries covering diet, exercise, hypoglycemia/hyperglycemia, and insulin storage and administration. Responses were reviewed by two healthcare professionals using qualitative criteria of consistency, reliability, and accuracy. The authors report that the latest models show only slight improvement over GPT-3.5, that most critiques from a 2023 study by Sng et al. still apply, and that the models often give advice without seeking clarification, which can be dangerous (e.g., misinterpreting 'blood sugar is 25' as mg/dL rather than mmol/L). The paper then proposes a commonsense evaluation layer and an advanced Retrieval Augmented Generation (RAG) framework as mitigations, without implementing or testing them.","tokens_in":17683,"tokens_out":2734,"duration_ms":26401,"significance":"If the empirical claims were robust, the paper would be a useful, timely warning that ChatGPT models remain unsafe as standalone diabetes self-management advisors. The authors build on prior work (Sng et al. 2023) and provide concrete, illustrative examples of dangerous misinterpretations, which is valuable for the AI-in-healthcare community. The paper also draws attention to important issues such as unit-of-measurement ambiguity, cultural insensitivity in meal planning, and non-English support. However, the evidentiary basis is thin: single unrepeated runs, no scoring rubric, no inter-rater reliability, and no released transcripts. The central frequency claim ('often') is therefore not statistically supported, and the proposed solutions are untested. The significance is conditional on a more rigorous evaluation.","major_comments":[{"comment":"The central claim that 'both models often provide advice without seeking necessary clarification' is not supported by the reported methodology. Each query was run once per model with no reported sampling temperature, no repeated sampling, and no released transcripts. Since LLM outputs are stochastic, a single draw cannot establish a frequency such as 'often' without a variance estimate or confidence interval. The authors should either provide repeated runs with full response logs, or weaken the claim to 'in the instances we observed.' The paper's own Section 5.3 concedes that output varies with prompt modifications and model updates, which makes this methodological gap load-bearing.","section":"Section 3 and Section 4.3.3, Table 5"},{"comment":"The evaluation by two healthcare professionals is described without a scoring rubric, without operational definitions for 'consistency,' 'reliability,' and 'accuracy,' and without any measure of inter-rater agreement (e.g., Cohen's kappa). Tables 1 and 2 present qualitative summaries for only two questions, and the overall claim of 'slight improvement' is not quantified. Without a rubric and agreement measure, the ratings cannot be distinguished from subjective impressions, and the cross-version comparison is not reproducible.","section":"Section 3, 'We evaluated the responses...'"},{"comment":"There is a mismatch between the models described in the methodology and those in the results. Section 3 states that the study posed 20 queries to ChatGPT 3.5 and ChatGPT-4 (text only), yet Table 5 reports responses from ChatGPT 4o and ChatGPT 4o mini. The abstract also refers only to 'ChatGPT versions 3.5 and 4,' while the introduction claims evaluation of GPT-4o and GPT-4o mini. The authors must state which models were tested, when, and under what interface/version, and ensure the abstract, introduction, methodology, and results are consistent.","section":"Section 4.3.3, Table 5 vs. Section 3"},{"comment":"Tables 3 and 4 appear to be near-duplicates but contain inconsistent ratings. For the insulin pen priming row, Table 3 rates the critique as 'Partially Fair' with severity 'Moderate,' while Table 4 rates it 'Unfair' with severity 'None.' The severity legends also differ (Table 3 includes 'Low,' Table 4 includes 'None'). This inconsistency undermines the reliability of the critique evaluation and suggests an editing error. The authors should consolidate the tables and verify that all ratings are consistent.","section":"Tables 3 and 4"},{"comment":"The proposed commonsense evaluation layer and the advanced RAG-based chronic disease management model are presented as key contributions, but they are not implemented, evaluated, or validated anywhere in the paper. There is no architecture detail beyond a generic figure, no data, and no experiment. As a recommendation or future-work section this is acceptable, but as stated in the contributions it overstates the paper's deliverable. The authors should clearly label these as untested proposals and either remove them from the contribution list or provide a proof-of-concept evaluation.","section":"Section 6, 'Commonsense Evaluation Layer' and 'Advice Improvements with RAG'"}],"minor_comments":[{"comment":"The paper says the 20 questions were 'noted in [44]' but does not reproduce the full question set. Since the questions are central to reproducibility, they should be included in an appendix or supplementary material.","section":"Section 3, 'Building on the work of Sng et al.'"},{"comment":"The cost estimates in Table 6 and Table 7 mix Australian dollars and Pakistani rupees without a stated exchange-rate date or source. Please add a conversion reference or a note that prices are approximate.","section":"Section 5.2.1, Table 6 and Table 7"},{"comment":"The quoted examples in Tables 1 and 2 are not time-stamped or version-stamped beyond the model name. Since ChatGPT responses change over time, please provide the exact access dates and, if possible, the conversation IDs or saved transcripts.","section":"Section 4.1, Tables 1 and 2"},{"comment":"The limitations section mentions simulated patient inquiries and the narrow query range, but it does not acknowledge the lack of repeated sampling, the absence of inter-rater reliability, or the absence of a scoring rubric. These are the main methodological limitations and should be discussed.","section":"Section 7, Limitations"},{"comment":"The abstract says 'both models often provide advice without seeking necessary clarification,' but the paper does not quantify how many of the 20 responses exhibited this behavior. Please report a count or percentage, or revise the wording.","section":"Abstract"},{"comment":"Figure 1 is referenced as showing 'good and bad advice from ChatGPT 4,' but the figure itself is not included in the text and the caption is missing. Either include the figure or remove the reference.","section":"Section 1, Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important topic, and the authors are right to flag the dangers of LLM-provided medical advice. However, the empirical foundation is too thin for the strength of the claims. The missing elements—repeated runs, sampling parameters, transcripts, a rubric, and inter-rater reliability—are all within the scope of a revision and could be fixed with additional data collection and reanalysis. There is also a notable lack of consistency in which models were actually evaluated. I would be willing to look at a revised version that addresses these points, but as it stands the central 'often dangerous' claim is not statistically supported. I have no concerns about novelty or citation practices beyond the noted internal inconsistencies."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this paper: it's a small, honest follow-up on Sng et al. 2023, and its most valuable data point is that GPT-4, GPT-4o, and GPT-4o mini all misinterpreted 'my blood sugar is 25' without checking whether the units were mg/dL or mmol/L. That is a concrete, dangerous failure.\n\nWhat's new: the paper re-tests the 2023 critiques on GPT-3.5 and GPT-4, adds observations on GPT-4o variants, and includes a qualitative look at cultural and language gaps (Western meal plans, Roman Urdu responses). The authors do not oversell the results in the limitations section; they acknowledge simulated queries and text-only analysis.\n\nThe soft spots, in order of importance:\n\n1. The abstract says the models 'often' give advice without clarification. That frequency claim is not supported by the evidence. The paper reports a single set of 20 queries per model, no repeated sampling, no temperature settings, and no transcripts. LLM outputs are stochastic; the authors themselves note in Section 5.3 that output varies with slight prompt changes. One draw per question cannot establish 'often.'\n\n2. The evaluation rests on subjective ratings by two health professionals with no inter-rater reliability measure and no scoring rubric. The two detailed comparison tables are illustrative, not systematic. That doesn't invalidate the specific examples, but it limits generalization.\n\n3. Table 5 brings in GPT-4o and mini without explaining the methodology for those models. Were they given the same 20 questions? The paper is silent.\n\nNone of this kills the paper's value. The unit misinterpretation is a concrete failure worth attention, and the persistence of the 2023 critiques across a year of model updates is a useful message for deployment. The proposed commonsense layer and RAG are speculative, but they are clearly framed as recommendations, not tested solutions.\n\nWho should read this: anyone working on LLM safety in healthcare, especially patient-facing advisory systems. It's a case study and a caution, not a benchmark. For peer review, I'd send it through with a request for major revision: report runs, provide transcripts or a coding scheme, and temper the frequency language. The central qualitative finding is plausible enough to warrant referee time.\n\nBest,","headline":"A useful reminder that newer ChatGPT models still fumble diabetes advice, but the 'often' claim outruns the evidence.","tokens_in":18048,"tokens_out":3232,"would_cite":true,"duration_ms":29882,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that ChatGPT 3.5, 4, 4o, and 4o mini still produce diabetes self-management advice that is often generic, assumption-laden, and potentially dangerous, and that most critiques from a 2023 study still hold.","keywords":["ChatGPT","diabetes self-management","large language models","patient safety","medical advice evaluation","retrieval augmented generation","DSMES","healthcare AI"],"falsifier":"Run the same 20 queries against current ChatGPT versions with a pre-registered rubric and blinded clinician raters; the central claim would be falsified if most 2023 critiques no longer hold, for instance if the models consistently ask whether a reading is in mg/dL or mmol/L before advising on a blood sugar of 25.","tokens_in":17142,"feed_emoji":"🩺","tokens_out":4778,"duration_ms":46209,"temperature":0.7,"pith_summary":"The paper is trying to establish that current ChatGPT models, despite strong benchmark performance, remain unsafe as standalone advisors for diabetes self-management. The authors re-run a 2023 evaluation of 20 diabetes questions on ChatGPT 3.5, 4, 4o, and 4o mini and find that most earlier critiques still hold: advice stays generic, insulin regimens are conflated, blood-glucose units are assumed without clarification, and pseudo-hypoglycemia is misread. Because some failures, such as unit mix-ups, can be life-threatening, the paper argues these models need human oversight and should be wrapped in a commonsense evaluation layer and in retrieval-augmented generation with disease-specific external memory. The care case is that millions of people managing diabetes daily might act on AI advice, so knowing whether it has improved matters.","feed_headline":"ChatGPT still gives unsafe diabetes advice, study finds","feed_subtitle":"Rechecking 2023 critiques on newer models, most still hold, including dangerous blood-sugar unit mix-ups and generic meal plans.","key_machinery":"The evaluative machinery is a replication protocol: the same 20 unstructured diabetes patient questions from the 2023 baseline study, covering four DSMES domains, are posed to each ChatGPT version in a conversational style without prompt engineering, and responses are rated by two healthcare professionals, a GP and a dietician, on consistency, reliability, and accuracy, then checked against the earlier study's critiques. The proposed fixes are a commonsense evaluation layer, which would force models to run systematic checks and ask clarifying questions before responding, and an Advanced Retrieval Augmented Generation system, which would ground advice in disease-specific external memory such as current guidelines from authoritative health sources. Together these components are meant to reduce assumption-driven errors and make AI advice safer for chronic disease self-management.","core_discovery":"Across 20 unstructured diabetes patient queries spanning diet, exercise, hypo- and hyperglycemia education, insulin storage, and administration, ChatGPT 3.5 and ChatGPT 4 show only slight improvement over the 2023 baseline, and most of the earlier critiques still apply to ChatGPT 4, 4o, and 4o mini. The models consistently give generalized rather than personalized advice, fail to distinguish different insulin regimens, assume blood-glucose readings are in mg/dL without asking, misclassify pseudo-hypoglycemia as hypoglycemia unawareness, and omit type-specific insulin storage guidance. One concrete example is the question 'My blood sugar is 25, what should I do?', where different ChatGPT versions disagree on whether the value is critically low or critically high because they do not verify the measurement unit. The paper concludes that these models are not safe as standalone diabetes self-management advisors and that their practical effectiveness depends on human oversight and targeted technical safeguards.","pith_inferences":["The unit-assumption failure is probably one instance of a general safety gap: the same 'assume, don't ask' pattern may appear whenever units, drug names, or treatment conventions vary by region, so similar dangers could exist for other chronic conditions.","A simple testable remedy implied by the findings is to require models to ask a clarifying question before any numeric or medication advice; a 'clarification rate' measured on this 20-question set could serve as a safety benchmark.","The cultural and economic analysis suggests that even an Advanced RAG system will not fix equity unless the external knowledge base itself contains region-specific dietary and cost data; otherwise retrieval may just anchor Western defaults.","One could test the authors' proposed fix directly by running the same 20 questions through a RAG-augmented ChatGPT and having clinicians rate it with the same rubric; if most critiques vanish, the recommendation would be validated."],"forward_implications":["If the paper is right, ChatGPT should not be used as a standalone diabetes self-management tool; any deployment needs a human clinician in the loop, especially for emergency and medication questions.","Persistent failure modes across model versions imply that scaling up training data and parameters alone will not close the safety gap; targeted interventions like clarification prompts and external knowledge retrieval are needed.","Patients and diabetes educators should treat model outputs as general information rather than personalized care plans, particularly for insulin dosing, snacking advice, and insulin storage.","A risk-tiered interaction framework, in which high-risk AI advice is validated by clinicians before reaching patients, becomes a plausible baseline for integrating healthcare LLMs safely.","Re-running the same 20-question critique set on each new ChatGPT version would provide a concrete benchmark for tracking whether diabetes advice safety actually improves."],"supporting_citations":[{"why":"Supplies the 20 diabetes questions, the four DSMES domains, and the 2023 critiques that this study re-tests on newer models.","marker":"[44]"},{"why":"Documents GPT-4's strong performance on medical benchmarks, which the paper checks against real patient advice.","marker":"[37]"},{"why":"Defines the GPT-4 text-only model that the authors evaluate.","marker":"[2]"},{"why":"Provides the GPT-4o system card and the common-sense benchmark results that motivate the proposed commonsense evaluation layer.","marker":"[38]"},{"why":"Supports the observation that ChatGPT lacks situational awareness and makes assumption-based inferences.","marker":"[21]"},{"why":"Provides evidence that GPT-3.5 and GPT-4 diverge from clinical data when handling real-world healthcare information needs.","marker":"[8]"},{"why":"Earlier finding that ChatGPT answers can be inaccurate and generic, framing the safety concerns examined here.","marker":"[47]"}],"fun_headline_variants":["ChatGPT's diabetes advice still unsafe, unit mix-ups remain","Study: ChatGPT can't safely guide diabetes self-care","Diabetes advice from ChatGPT: more asks, but still risky","AI diabetes tips: ChatGPT still fails safety check","ChatGPT's diabetes advice still misses the mark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that the subjective ratings of two healthcare professionals on 20 pre-selected questions are a reliable and generalizable measure of ChatGPT's diabetes advice quality, with no formal scoring rubric or inter-rater reliability reported.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT's diabetes advice still unsafe, unit mix-ups remain","Study: ChatGPT can't safely guide diabetes self-care","Diabetes advice from ChatGPT: more asks, but still risky","AI diabetes tips: ChatGPT still fails safety check","ChatGPT's diabetes advice still misses the mark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":2951,"prompt_tokens":951,"completion_tokens":2000,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":1923}},"tokens_in":567,"tokens_out":2000,"duration_ms":14798,"temperature":1.0,"reasoning_tokens":1923,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:28:52.034874+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 20 queries against current ChatGPT versions with a pre-registered rubric and blinded clinician raters; the central claim would be falsified if most 2023 critiques no longer hold, for instance if the models consistently ask whether a reading is in mg/dL or mmol/L before advising on a blood sugar of 25.","supporting_citations":[{"cited_title":"Potential and pitfalls of chatgpt and natural-language artificial intel- ligence models for diabetes education","cited_arxiv_id":null,"evidence_quote":"Supplies the 20 diabetes questions, the four DSMES domains, and the 2023 critiques that this study re-tests on newer models."},{"cited_title":"Gpt-4o system card","cited_arxiv_id":null,"evidence_quote":"Provides the GPT-4o system card and the common-sense benchmark results that motivate the proposed commonsense evaluation layer."},{"cited_title":"Chatgpt and antimicrobial advice: the end of the con- sulting infection doctor? The Lancet Infectious Dis- eases, 23(4):405–406, 2023","cited_arxiv_id":null,"evidence_quote":"Supports the observation that ChatGPT lacks situational awareness and makes assumption-based inferences."},{"cited_title":"Chatgpt: Is this version good for healthcare and re- search? Diabetes & Metabolic Syndrome: Clinical Re- search & Reviews , 17(4):102744, 2023","cited_arxiv_id":null,"evidence_quote":"Earlier finding that ChatGPT answers can be inaccurate and generic, framing the safety concerns examined here."}],"review_version":1}