REVIEW 4 major objections 4 minor 2 cited by
AI-based Clinical Decision Support for Primary Care: A Real-World Study
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read In live primary care clinics, an LLM-based safety net cut physician-rated diagnostic errors by 16% and treatment errors by 13%.
desk verdict First large real-world LLM CDS study with careful statistics, but the primary endpoint measures documentation quality that the tool directly coaches, so the headline error reduction is likely inflated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is AI Consult, an LLM-based clinical 'safety net' embedded in the electronic medical record. The load-bearing mechanism is the red/yellow/green severity interface, where the model reviews each visit's documentation at asynchronous workflow decision points (vitals, clinical notes, investigations, diagnosis, treatment) and returns a severity color with a rationale and an action. The design includes a shadow mode that logs would-have-been alerts for the control group, and a 'left in red' metric tracking visits where the final alert remains red, which drives active deployment coaching.
What would settle it
A reanalysis that controlled for clinical-note length and documentation completeness, or a randomised trial using adjudicated patient outcomes such as hospitalisation, return visits, or delayed diagnosis rather than documentation-based ratings, would settle whether the error reduction is real. Concretely, if the relative risk reduction for treatment errors dropped to near zero once note length and documentation detail were included as covariates, the central claim would be refuted.
Extended reading notes
Core claim
The central discovery, stated on the paper's own terms, is that LLM-based clinical decision support deployed as an asynchronous safety net in live primary care produces measurable reductions in physician-rated clinical errors. The claims are a 31.8% relative reduction in history-taking errors, 10.3% in investigation errors, 16.0% in diagnostic errors, and 12.7% in treatment errors when comparing visits of clinicians with AI Consult to those without, with the strongest effects after an induction period of active deployment. The authors further claim that the red/yellow/green alert severities correlate with physician-rated quality, that clinicians learn to avoid errors even before alerts fire, and that no patient safety reports attributed harm to the tool's advice.
Load-bearing premise
The paper's central claim rests on the assumption that independent physicians' ratings of the clinical documentation are an unbiased measure of true clinical error, which is strained because the AI-group clinicians wrote longer notes, the physician raters agreed with each other only fairly (Fleiss kappa 0.22 to 0.29), and there was no pre-rollout baseline to anchor the comparison.
Editorial extensions
If this is right
- If replicated, LLM-based clinical decision support can function as a real-time safety net in high-volume primary care without requiring clinicians to request assistance.
- The error reductions imply that a capable general-purpose model, prompted with local guidelines and context, can be sufficient; no fine-tuned or specialised clinical model is required.
- The learning effect (fewer visits starting with red alerts over time) suggests the tool functions as case-based continuing education, potentially improving care even after alerts stop firing.
- The annual projection of roughly 22,000 fewer diagnostic errors and 29,000 fewer treatment errors at the study network alone indicates the scale of potential benefit in similar high-volume settings.
- The absence of a patient-reported outcome difference and the added clinician visit time flag a trade-off: the tool improved documented clinical quality while costing time, and the downstream patient benefit remains to be measured.
Reading between the lines
- Because the outcome is physician-rated documentation quality and clinicians with the tool wrote longer notes, some of the measured error reduction may reflect improved documentation rather than improved clinical decisions; a study with chart-independent outcomes would separate these.
- The learning effect implies that the tool's benefit may grow with continued use as clinicians internalise alert patterns, so a longer deployment might show larger effects than the 10-week study observed.
- The LLM raters found larger differences than physician raters; if model-based rating is validated on a subset, it could make routine quality monitoring scalable, but its use of similar LLM-based logic could overstate improvement for LLM-assisted documentation.
- The finding that both recorded deaths were judged potentially preventable if alerts had been heeded suggests the bottleneck may shift from model capability to clinician trust and adherence, which is an implication for how such tools are deployed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a pragmatic, cluster-assigned quality-improvement study of an LLM-based clinical decision support tool ("AI Consult") deployed at Penda Health, a primary-care network in Nairobi, Kenya. Across 39,849 consenting patient visits from 15 clinics, half of the clinicians in each clinic were randomized to have AI Consult available while the other half did not, with shadow mode recording what alerts would have fired for the non-AI group. The primary outcome is physician-rated documentation quality: independent, blinded physicians assigned Likert 1-5 scores for history-taking, investigations, diagnosis, and treatment, with scores 1-2 defined as clinically meaningful errors. The paper reports relative reductions of 31.8% for history errors, 10.3% for investigation errors, 16.0% for diagnostic errors, and 12.7% for treatment errors, robust to GEE and modified Poisson adjustments, and projects tens of thousands of annual averted errors at Penda if the tool were scaled. Secondary analyses examine failure modes, clinician uptake via a "left in red" metric, a claimed learning effect from declining "started red" rates, clinician surveys, and patient-reported outcomes, which showed no significant differences. The authors conclude that careful implementation and active deployment, not model capability alone, were critical to the observed benefits.
Significance. If the central claim is accepted, this is an important early demonstration that LLM-based decision support can reduce errors in live primary care, and the study has notable methodological strengths: intent-to-treat analysis, cluster-robust GEE and modified Poisson models, blinded and international physician raters, Benjamini-Hochberg correction, a large sample, and public release of analysis code. The shadow-mode design for the non-AI group is a valuable feature that permits comparison of would-have-been alerts. However, the primary endpoint is physician-rated documentation quality, and the intervention is explicitly designed to improve documentation completeness. The reported error reductions may therefore partly reflect improved charting rather than improved clinical decisions. The fair inter-rater reliability (Fleiss kappa 0.22-0.29), the longer notes in the AI group, and the absence of a pre-rollout baseline make this measurement-validity threat material. The learning-effect analysis is also circular in using the model's own red alerts as both the intervention signal and the outcome.
major comments (4)
- [Section 3.8, Table 19, Fig. 26] The primary outcome is physician-rated documentation quality, not a direct measure of clinical decisions, and the intervention is designed to improve documentation completeness. The prompts explicitly ask clinicians to record MUAC, respiratory rate, missing history details, and additional diagnoses (Appendix E), and the error definitions include "key details in the history are missing," "key investigations are missing," and "clinically relevant additional diagnosis is missing" (Table 19). The paper also reports that AI-group notes were longer (600 vs 450 characters, Fig. 26). Because raters see only the documentation, longer and more complete notes can mechanically lower Likert 1/2 rates even if the clinician's actual diagnostic and treatment decisions are unchanged. The fair inter-rater agreement (Fleiss kappa 0.22-0.29) and the absence of a pre-rollout baseline make this alternative explanation difficult to exclude. Please add sensitivity analyses that adjust for note length or documentation completeness, and ideally report at least one outcome less dependent on documentation, such as medication orders, laboratory orders, referrals, or the subset of ratings where the rater explicitly identified a wrong decision rather than a missing note.
- [Section 3.3, Table 1] The randomized comparison has no pre-rollout measurement of the outcome. Clinician-level randomization within clinics is a reasonable design, but only 15 clinics were included, 57 AI and 49 non-AI clinicians contributed, and the arms show baseline imbalance in clinic region (42.8% vs 34.2% in Thika Road Corridor). The GEE and modified Poisson models adjust for clinic and patient covariates, but they cannot adjust for unmeasured clinician-level differences in documentation habits or baseline error rates. I recommend reporting a baseline or pre-period analysis (e.g., from shadow-mode data or from a pre-study chart sample) to show that the two groups had comparable documentation quality before the intervention, or explicitly acknowledging that the reported effect could partly reflect clinician-level confounding.
- [Section 4.1, Fig. 6b, Discussion] The claimed learning effect is circular. The "started red" rate is computed from AI Consult's own red/yellow/green classifications, which are generated by the same model whose alerts are the intervention. A decline in this rate among AI-group clinicians could reflect changes in documentation style that the model no longer flags, rather than genuine improvement in clinical reasoning. The non-AI group's "started red" rate is computed from shadow-mode calls, but prompt and threshold changes during the induction period (documented in Section 3.3) make the metric unstable over time. Please corroborate the learning claim with physician-rated error rates stratified by study week or with a pre-specified test of whether the decline in started-red rate tracks an independent clinical outcome.
- [Section 4.1, Table 3] The extrapolation to "22,000 fewer diagnostic errors and 29,000 fewer treatment errors annually at Penda" is not supported by the measurement. The outcomes are physician-rated documentation errors, not confirmed clinical errors, and the absolute numbers apply the relative risk reduction to all 400,000 annual visits without accounting for the fact that the primary endpoint is a rating construct with fair inter-rater reliability. The NNTs likewise quantify the number of visits needed to avoid one rated documentation error, not one patient harm event. Please reframe these projections as "rated errors" and add a caveat that they inherit the uncertainty of the rating-based outcome.
minor comments (4)
- [Figure 26] The text in Section 4.2 refers to Fig. 26 as showing clinical note length over time, but the caption in Appendix D.7 describes it as showing the rate of clinician thumbs-up feedback on AI Consult responses. Please correct the figure number or caption.
- [Section 3.8, footnote 5] The handling of missing structured chief-complaint data after April 9 is described in a footnote, but it is not clear whether the history RRR in Table 3 is restricted to visits before April 10 or how many visits were excluded. Please state this explicitly in the main text or in the Table 3 notes.
- [Section 4.2] The term "attending time" is used without an explicit definition; please clarify whether it is the total visit duration, clinician idle time, or another measure, and state how it is recorded in the EMR.
- [References] The KLAS reference lists the year as 2003 while the URL and text suggest a 2023 report; please verify and correct the citation year.
Circularity Check
Primary endpoint uses independent physician raters, but two secondary analyses are self-referential: the learning effect is measured by AI Consult's own red flags, and OpenAI LLMs rate documentation coached by an OpenAI LLM.
-
other
[Section 4.1, 'Clinicians in the AI group learned to avoid common mistakes over time' (Fig. 6b)]
"We also examine the proportion of visits where AI Consult started red–that is, where the first AI call for any category was red. In the AI group, this rate drops from 45% at the start of the study to 35% at the end of the study, while staying steady at 45-50% in the non-AI group during the study (Fig. 6b). This suggests that AI Consult is training clinicians to avoid common mistakes even prior to AI Consult alerts."
The 'learning effect' claim uses AI Consult's own red-flag classifications as the outcome metric. The intervention and the measurement are produced by the same model: a declining started-red rate means the model's alerts changed, not independently that clinical errors changed. Without an external anchor to patient outcomes or physician judgments, the decrease in the AI group could reflect model calibration drift, prompt iteration during the induction period, or clinicians learning to satisfy the model's documentation preferences rather than improving care. This is a secondary analysis and does not drive the main physician-rated endpoint.
-
other
[Section 3.8 'LLM rater analysis' and Section 5.1 Limitations]
"While greater effect sizes may be the result of Goodhart's Law (clinician documentation is assessed by an LLM in AI Consult as well), the greater model-physician agreement compared to physician-physician agreement suggests that LLM ratings, if validated via physician agreement on a subset of cases, may be a way to scale up both routine quality improvement and studies like this one."
The LLM-rater analysis presents GPT-4.1 and o3 effect sizes as robustness for the main finding, but these raters are OpenAI LLMs grading documentation that a different OpenAI LLM (GPT-4o in AI Consult) actively coached. The paper itself concedes the larger effect sizes may reflect Goodhart's law: the documentation was shaped to satisfy an LLM grader, so an LLM grader is not an independent adjudicator. This analysis is secondary; the primary claim rests on the 108-physician panel.
full rationale
The paper's central comparison (16% fewer diagnostic errors and 13% fewer treatment errors) rests on ratings by 108 independent physicians blinded to group assignment, with golden examples, training, and dual rating of about 25% of visits. That endpoint is not AI Consult's own output, so the main claim is not circular by construction. The skeptical concern that AI Consult coaches documentation completeness—and that Likert 1/2 error definitions include missing history details, missing investigations, and missing additional diagnoses—is a genuine measurement-validity threat, but it is a construct-validity concern rather than a logical reduction of the prediction to the model's inputs; I therefore do not count it as formal circularity. Two secondary analyses are self-referential: the 'learning effect' uses AI Consult's own started-red rate as the outcome, and the LLM-rater analysis uses OpenAI LLMs to grade documentation that an OpenAI LLM helped produce, with the paper itself noting Goodhart's law as a possible explanation for the larger effect sizes. These are exploratory and do not drive the primary endpoint, so the overall circularity score is 2.
Assumptions & free parameters
free parameters (1)
- Red/yellow/green severity thresholds =
not reported (tuned)
assumptions (4)
- domain assumption Physician chart-review Likert ratings are a valid measure of clinical error.
- domain assumption Clinician-level randomization within clinics produced exchangeable groups.
- domain assumption No substantial contamination between AI and non-AI clinicians in the same clinic.
- standard math GEE and modified Poisson models are correctly specified.
Cite this review
Pith. "Pith review of AI-based Clinical Decision Support for Primary Care: A Real-World Study." pith.science (2026). https://pith.science/paper/5BZKISFN
@misc{pith2026250716947,
author = {Pith},
title = {Pith review of: AI-based Clinical Decision Support for Primary Care: A Real-World Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/5BZKISFN}},
note = {Machine review of arXiv:2507.16947}
}
read the original abstract
We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya, we studied AI Consult, a tool that serves as a safety net for clinicians by identifying potential documentation and clinical decision-making errors. AI Consult integrates into clinician workflows, activating only when needed and preserving clinician autonomy. We conducted a quality improvement study, comparing outcomes for 39,849 patient visits performed by clinicians with or without access to AI Consult across 15 clinics. Visits were rated by independent physicians to identify clinical errors. Clinicians with access to AI Consult made relatively fewer errors: 16% fewer diagnostic errors and 13% fewer treatment errors. In absolute terms, the introduction of AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda alone. In a survey of clinicians with AI Consult, all clinicians said that AI Consult improved the quality of care they delivered, with 75% saying the effect was "substantial". These results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake. We hope this study demonstrates the potential for LLM-based clinical decision support tools to reduce errors in real-world settings and provides a practical framework for advancing responsible adoption.
Figures
Figures from the paper (24 more)
Forward citations
Cited by 2 Pith papers
-
A global log for medical AI
MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.
-
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
Passing medical exams does not make LLMs safe for autonomous triage: they fail to seek missing red flags, and current benchmarks do not test that.
Reference graph
Works this paper leans on
-
[1]
Evaluate the patient’s chief complaint and vital signs (and MUAC for children ages 6 months–5 years)
-
[2]
Determine whether there are urgent or concerning findings that may indicate a medical emergency (Red), incomplete or suboptimal documentation or potential concerns (Yellow), or if everything is appropriate and non-urgent (Green)
-
[3]
Provide concise, actionable recommendations to improve patient safety and care quality. Severity Thresholds
-
[4]
patient is well-appearing, normal affect
Actionable Recommendations •If Red: Offer urgent steps (e.g., re-check vitals, immediate advanced care for suspected emergencies). •If Yellow: Suggest needed clarifications or missing vitals. •If Green: Encourage routine next steps; no critical gaps. Output Structure You must return exactly one severity level in JSON, with an explanatory Reason and an Act...
-
[5]
Medications are present but inappropriate
Likely inappropriate class of antibiotics used (e.g., amox/clavulanic acid used when amoxicillin is appropriate) *if you select this choice, you should also select choice “Medications are present but inappropriate”
-
[6]
Referrals are present but inappropriate 8
Referrals are missing 7. Referrals are present but inappropriate 8. Needed procedures are missing 9. Procedures are present but inappropriate 10. Needed escalations of care are missing 11. Escalations of care are present but inappropriate 12. None of the above Additional resources for physicians Common local brand name pharmaceutical list that physicians ...
-
[7]
Chief complaint of severe headache with BP 180/110 mmHg—possible hypertensive emergency
Red •Potential emergency based on the chief complaint and abnormal vitals (e.g., severe chest pain + very high BP, or severe headache + hypertensive crisis). •All vitals are missing (critical omission). •If a pregnant patient’s complaint and vitals suggest a severe complication (e.g., very high BP, severe edema, etc.). •Example: “Chief complaint of severe...
-
[8]
•Some essential vitals are missing but not all
Yellow •Concerning chief complaint (e.g., chest pain) but vitals do not clearly indicate an emergency; additional assessment is needed. •Some essential vitals are missing but not all . •Respiratory complaints without documented SpO2 or Respiratory Rate . •If the MUAC or other vital sign is borderline, or mild abnormalities that need follow-up but are not emergent
Show all 48 references
-
[9]
Vitals within normal limits, mild sore throat, no red flags
Green •All relevant vitals are documented, no signs of emergent danger in the chief complaint or vitals. •Example: “Vitals within normal limits, mild sore throat, no red flags.” Key Principles
-
[10]
•Children under 12 : Temperature, Pulse, (BP is not expected), and for ages 6 months–5 years, MUAC is recommended but not mandatory 80 •Pregnant Patients : BP is crucial
Essential Vitals •Adults : Temperature, Pulse (HR), Blood Pressure (BP), Height, Weight, and Calculated BMI, and, if respiratory complaints, recommend SpO2. •Children under 12 : Temperature, Pulse, (BP is not expected), and for ages 6 months–5 years, MUAC is recommended but no...
-
[11]
MUAC Interpretation •Red : Severe malnutrition (urgent) •Yellow : Moderate malnutrition •Green : No malnutrition •Do not request MUAC for ages outside 6 months–5 years unless specifically indicated
-
[12]
Respiratory Rate •While helpful, respiratory rate is not critical for all patients (except those with respiratory complaints, in which case missing RR or SpO2 triggers Red)
-
[14]
•No critical tests are missing; no irrelevant or unjustified tests are ordered
Green •The investigations ordered are appropriate and comprehensive for the clinical scenario. •No critical tests are missing; no irrelevant or unjustified tests are ordered. •Example: A strep test ordered for a patient with sore throat and exudative tonsillitis, or a urine di...
-
[15]
•OR there is at least one low-value or marginally justified test ordered
Yellow •Some recommended investigations are missing or questionable based on the history / exam, but not so critical as to seriously endanger the patient. •OR there is at least one low-value or marginally justified test ordered. •Example: Mild pallor noted but no full haemogra...
-
[16]
Add test X,
Red •Essential diagnostic investigations are omitted, posing a risk of delayed or inaccurate diagnosis. •Clearly inappropriate tests are ordered, showing a major mismatch with the documented presentation. •Example: A patient with severe chest pain but no cardiac or respiratory...
-
[17]
Evaluate the clinician’s diagnosis against the visit’s documentation (patient history, exam findings, vitals, labs, etc.)
-
[18]
Assess if the diagnosis is appropriate, missing, incomplete, or incorrectly severe given the local epidemi- ology and available resources
-
[19]
Severity Thresholds
Provide concise, actionable recommendations to guide safe and quality patient care. Severity Thresholds
-
[20]
•No significant mismatch with history, vitals, labs, or local context
Green •The listed diagnosis (or diagnoses) accurately reflects the clinical documentation. •No significant mismatch with history, vitals, labs, or local context. •The clinician may safely proceed with management of these diagnoses. •Example: If the patient presents with dysuri...
-
[21]
– Additional testing or more thorough documentation is advisable before finalizing
Yellow •The listed diagnosis broadly aligns with the documentation, but: – There is some uncertainty or missing details preventing a definitive conclusion (e.g., possible severe pathol- ogy but not fully confirmed). – Additional testing or more thorough documentation is advisa...
-
[22]
simple cystitis
Red •A serious mismatch: The listed diagnosis is incompatible with the clinical findings, or a critical diagnosis is missing. •Could result in dangerous consequences if not corrected. •Severe diagnoses listed are not supported by the presentation, or a severe condition is clea...
-
[23]
She has had these symptoms before and was diagnosed with UTI
Green Example Age: 25y Gender: Female Vitals: Temperature: 37.80°C Pulse: 80 bpm Blood Pressure: 120/78 Respiratory Rate: 18 SPO2: 99 Chief Complaint: Dysuria, urinary frequency Clinical notes: Pt complains of dysuria and urinary frequency x2 days. She has had these symptoms b...
-
[24]
Yellow Example Clinical Documentation: Age: 16y Gender: Male Vitals: Temperature: 38.50°C Pulse: 90 bpm Blood Pressure: 110/70 Respiratory Rate: 20 SPO2: 98 Chief Complaint: Right lower quadrant abdominal pain, mild nausea Physical Exam: Mild tenderness in RLQ but no rebound o...
-
[25]
Red Example 90 Clinical Documentation Age: 35y Gender: Female Vitals: Temperature: 39.20°C Pulse: 105 bpm Blood Pressure: 130/85 Respiratory Rate: 22 SPO2: 98 Chief Complaint: Flank pain, fever, nausea Physical Exam: Notable costovertebral angle tenderness Lab Results: WBC cou...
-
[26]
Evaluate the clinician’s treatment plan against the visit documentation (vitals, diagnosis, labs, etc.)
-
[27]
Identify if the treatment is safe, evidence-based, and aligned with local guidelines (e.g., MoH Kenya, IMNCI/WHO)
-
[28]
Severity Thresholds
Provide concise, actionable recommendations to ensure appropriate and safe patient care. Severity Thresholds
-
[29]
Red •Serious mismatch between treatment and diagnosis. •Unsafe or unnecessary medications (e.g., antibiotics for a confirmed viral illness, sedating antihistamines in young children, monteleukast for respiratory infections without asthma). •Omission of essential medications wh...
-
[30]
•Some prescriptions listed are of dubious value to the patient (e.g., cough syrups)
Yellow •Treatment plan mostly aligns with the documented diagnosis, but: •Minor adjustments to dosage/duration are recommended, or while the medication choice is acceptable, it is not considered a first-line treatment for the condition. •Some prescriptions listed are of dubiou...
-
[31]
•No critical omissions or unnecessary interventions
Green •Treatment plan is complete, accurate, and in compliance with relevant guidelines. •No critical omissions or unnecessary interventions. Specific Guidelines to note
-
[32]
Key IMNCI/WHO Guidance for Dehydration in Children ¡5 Years
-
[33]
•Then 70 mL/kg over 2.5 hours (¿ 12 months) or 5 hours (¡ 12 months)
Severe Dehydration •IV Ringer’s Lactate at 30 mL/kg over 30 min (if child ¿ 12 months) or 60 min (¡ 12 months). •Then 70 mL/kg over 2.5 hours (¿ 12 months) or 5 hours (¡ 12 months). •If IV access is not possible, ORS via nasogastric tube at 120 mL/kg over 6 hours
-
[34]
Some Dehydration •Oral Rehydration Solution (ORS) at 75 mL/kg over 4 hours. 92
-
[35]
All cases should receive zinc supplementation
No Dehydration •ORS 10 mL/kg after each loose stool. All cases should receive zinc supplementation
-
[36]
Septrin (cotrimoxazole) is not a recommended first-line treatment due to its use in TB management
Urinary Tract Infection Management In Kenya, Nitrofurantoin and Cephalosporins are appropriate first-line therapy for management of UTI in adults and pregnant women. Septrin (cotrimoxazole) is not a recommended first-line treatment due to its use in TB management
-
[37]
5”, say “so you’re feeling much better
Note that Zefcolin (brand name) is a cough syrup and not a cephalosporin antibiotic; it can be used to relieve cough symptoms associated with upper respiratory tract infection in adults and children over 2 years. Output Structure You must return exactly one severity level in J...
-
[38]
∗Please make sure to flag severe outcomes like hospital admission, ICU admission, or death here
If they do, confirm their answer verbally! – Ask this question even if the patient is feeling better!They may be feeling better because they have already gone to another hospital or chemist 95 – If the patient plans to visit another clinic but hasn’t yet, answer “No”!This ques...
-
[39]
Key details in the history are missing (e.g., characterization of chief complaint is lacking key elements such as onset, duration, or associated symptoms, etc
Chief complaint is absent 2. Key details in the history are missing (e.g., characterization of chief complaint is lacking key elements such as onset, duration, or associated symptoms, etc. or pertinent medical history such as travel/sexual/family history are missing when they ...
-
[40]
Documentation of relevant systems on physical exam are absent (e.g., respiratory exam in a patient with cough, description of the rash in a patient with skin findings)
-
[41]
Blood group – self request
Pertinent vital signs are absent 5. None of the above Investigations form and questions Task Specific Clinical Note Investigations: Likert Score : Grade the investigations ordered (or lack of them) on appropriateness, given the clinical documentation. Keep in mind that not all...
-
[42]
Unjustified investigations are ordered
-
[43]
elevated blood pressure reading
None of the above (Any investigations ordered are indicated and no key investigations are missing) Diagnosis form and questions Task Specific Clinical Note Diagnosis Likert Score: Grade whether the diagnosis (including primary diagnosis, any additional diagnoses, or the listed...
-
[44]
allergic rhinitis
Primary diagnosis is too specific to be supported based on current documentation or investigations (e.g., using “allergic rhinitis” as the diagnosis rather than “rhinitis”, where it’s clear that rhinitis is present but documentation does not support whether it is a viral, bact...
-
[45]
Clinically relevant additional diagnosis is missing (e.g
Additional diagnosis is likely incorrect 5. Clinically relevant additional diagnosis is missing (e.g. malnutrition) 6. None of the above (All diagnoses are likely correct and no clinically relevant diagnoses are missing) Treatment form and questions Task Specific Clinical Note...
-
[46]
Medications are present but inappropriate
Likely inappropriate use of antibiotics overall (e.g., antibiotics are given for a likely viral infection) *if you select this choice, you should also select choice “Medications are present but inappropriate”
-
[2006]
doi: 10.1186/1748-5908-1-1
ISSN 1748-5908. doi: 10.1186/1748-5908-1-1. 29 S. L. Fleming, A. Lozano, W. J. Haberkorn, J. A. Jindal, E. Reis, R. Thapa, L. Blankemeier, J. Z. Genkins, E. Steinberg, A. Nayak, et al. Medalign: A clinician-generated dataset for instruction following with electronic medical re...
-
[2018]
doi: 10.1001/jama.2017.18391. S. Bedi, H. Cui, M. Fuentes, A. Unell, M. Wornow, J. M. Banda, N. Kotecha, T. Keyes, Y. Mai, M. Oez, et al. Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025. M. Benary, X. D. Wang, M. Sc...
2017
-
[2020]
elevated blood pressure reading
URLhttps://www.ajol.info/index.php/eamj/article/view/205283. D. McDuff, M. Schaekermann, T. Tu, and et al. Towards accurate differential diagnosis with large language models.Nature, 626:102–118, 2025. doi: 10.1038/s41586-025-08869-4. B. Middleton, D. F. Sittig, and A. Wright. ...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.