Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

AI-based Clinical Decision Support for Primary Care: A Real-World Study

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read In live primary care clinics, an LLM-based safety net cut physician-rated diagnostic errors by 16% and treatment errors by 13%.

desk verdict First large real-world LLM CDS study with careful statistics, but the primary endpoint measures documentation quality that the tool directly coaches, so the headline error reduction is likely inflated. read the letter →

arxiv 2507.16947 v1 pith:5BZKISFN submitted 2025-07-22 cs.CL

classification cs.CL
keywords LLMclinicaldecisionsupportprimarycarereal-worlddeploymenterrorreductiontraffic-lightalertslearningeffectKenyapatientsafety
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a quality-improvement study of AI Consult, a large language model-based decision-support tool embedded in the electronic medical record at a network of primary care clinics in Nairobi, Kenya. Across 39,849 patient visits, clinicians with access to the tool had fewer physician-rated clinical errors than those without: 16% fewer diagnostic errors and 13% fewer treatment errors in the main study period, with larger relative reductions for history-taking. The study argues that a model-as-safety-net design, tuned to clinical workflow and paired with active deployment, can reduce errors in real-world primary care rather than only on benchmarks. The paper also reports a learning effect, where clinicians with the tool increasingly avoided triggering the tool's red alerts over time.

What carries the argument

The central object is AI Consult, an LLM-based clinical 'safety net' embedded in the electronic medical record. The load-bearing mechanism is the red/yellow/green severity interface, where the model reviews each visit's documentation at asynchronous workflow decision points (vitals, clinical notes, investigations, diagnosis, treatment) and returns a severity color with a rationale and an action. The design includes a shadow mode that logs would-have-been alerts for the control group, and a 'left in red' metric tracking visits where the final alert remains red, which drives active deployment coaching.

What would settle it

A reanalysis that controlled for clinical-note length and documentation completeness, or a randomised trial using adjudicated patient outcomes such as hospitalisation, return visits, or delayed diagnosis rather than documentation-based ratings, would settle whether the error reduction is real. Concretely, if the relative risk reduction for treatment errors dropped to near zero once note length and documentation detail were included as covariates, the central claim would be refuted.

Watch

Extended reading notes

Core claim

The central discovery, stated on the paper's own terms, is that LLM-based clinical decision support deployed as an asynchronous safety net in live primary care produces measurable reductions in physician-rated clinical errors. The claims are a 31.8% relative reduction in history-taking errors, 10.3% in investigation errors, 16.0% in diagnostic errors, and 12.7% in treatment errors when comparing visits of clinicians with AI Consult to those without, with the strongest effects after an induction period of active deployment. The authors further claim that the red/yellow/green alert severities correlate with physician-rated quality, that clinicians learn to avoid errors even before alerts fire, and that no patient safety reports attributed harm to the tool's advice.

Load-bearing premise

The paper's central claim rests on the assumption that independent physicians' ratings of the clinical documentation are an unbiased measure of true clinical error, which is strained because the AI-group clinicians wrote longer notes, the physician raters agreed with each other only fairly (Fleiss kappa 0.22 to 0.29), and there was no pre-rollout baseline to anchor the comparison.

Editorial extensions

If this is right

  • If replicated, LLM-based clinical decision support can function as a real-time safety net in high-volume primary care without requiring clinicians to request assistance.
  • The error reductions imply that a capable general-purpose model, prompted with local guidelines and context, can be sufficient; no fine-tuned or specialised clinical model is required.
  • The learning effect (fewer visits starting with red alerts over time) suggests the tool functions as case-based continuing education, potentially improving care even after alerts stop firing.
  • The annual projection of roughly 22,000 fewer diagnostic errors and 29,000 fewer treatment errors at the study network alone indicates the scale of potential benefit in similar high-volume settings.
  • The absence of a patient-reported outcome difference and the added clinician visit time flag a trade-off: the tool improved documented clinical quality while costing time, and the downstream patient benefit remains to be measured.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the outcome is physician-rated documentation quality and clinicians with the tool wrote longer notes, some of the measured error reduction may reflect improved documentation rather than improved clinical decisions; a study with chart-independent outcomes would separate these.
  • The learning effect implies that the tool's benefit may grow with continued use as clinicians internalise alert patterns, so a longer deployment might show larger effects than the 10-week study observed.
  • The LLM raters found larger differences than physician raters; if model-based rating is validated on a subset, it could make routine quality monitoring scalable, but its use of similar LLM-based logic could overstate improvement for LLM-assisted documentation.
  • The finding that both recorded deaths were judged potentially preventable if alerts had been heeded suggests the bottleneck may shift from model capability to clinician trust and adherence, which is an implication for how such tools are deployed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript reports a pragmatic, cluster-assigned quality-improvement study of an LLM-based clinical decision support tool ("AI Consult") deployed at Penda Health, a primary-care network in Nairobi, Kenya. Across 39,849 consenting patient visits from 15 clinics, half of the clinicians in each clinic were randomized to have AI Consult available while the other half did not, with shadow mode recording what alerts would have fired for the non-AI group. The primary outcome is physician-rated documentation quality: independent, blinded physicians assigned Likert 1-5 scores for history-taking, investigations, diagnosis, and treatment, with scores 1-2 defined as clinically meaningful errors. The paper reports relative reductions of 31.8% for history errors, 10.3% for investigation errors, 16.0% for diagnostic errors, and 12.7% for treatment errors, robust to GEE and modified Poisson adjustments, and projects tens of thousands of annual averted errors at Penda if the tool were scaled. Secondary analyses examine failure modes, clinician uptake via a "left in red" metric, a claimed learning effect from declining "started red" rates, clinician surveys, and patient-reported outcomes, which showed no significant differences. The authors conclude that careful implementation and active deployment, not model capability alone, were critical to the observed benefits.

Significance. If the central claim is accepted, this is an important early demonstration that LLM-based decision support can reduce errors in live primary care, and the study has notable methodological strengths: intent-to-treat analysis, cluster-robust GEE and modified Poisson models, blinded and international physician raters, Benjamini-Hochberg correction, a large sample, and public release of analysis code. The shadow-mode design for the non-AI group is a valuable feature that permits comparison of would-have-been alerts. However, the primary endpoint is physician-rated documentation quality, and the intervention is explicitly designed to improve documentation completeness. The reported error reductions may therefore partly reflect improved charting rather than improved clinical decisions. The fair inter-rater reliability (Fleiss kappa 0.22-0.29), the longer notes in the AI group, and the absence of a pre-rollout baseline make this measurement-validity threat material. The learning-effect analysis is also circular in using the model's own red alerts as both the intervention signal and the outcome.

major comments (4)
  1. [Section 3.8, Table 19, Fig. 26] The primary outcome is physician-rated documentation quality, not a direct measure of clinical decisions, and the intervention is designed to improve documentation completeness. The prompts explicitly ask clinicians to record MUAC, respiratory rate, missing history details, and additional diagnoses (Appendix E), and the error definitions include "key details in the history are missing," "key investigations are missing," and "clinically relevant additional diagnosis is missing" (Table 19). The paper also reports that AI-group notes were longer (600 vs 450 characters, Fig. 26). Because raters see only the documentation, longer and more complete notes can mechanically lower Likert 1/2 rates even if the clinician's actual diagnostic and treatment decisions are unchanged. The fair inter-rater agreement (Fleiss kappa 0.22-0.29) and the absence of a pre-rollout baseline make this alternative explanation difficult to exclude. Please add sensitivity analyses that adjust for note length or documentation completeness, and ideally report at least one outcome less dependent on documentation, such as medication orders, laboratory orders, referrals, or the subset of ratings where the rater explicitly identified a wrong decision rather than a missing note.
  2. [Section 3.3, Table 1] The randomized comparison has no pre-rollout measurement of the outcome. Clinician-level randomization within clinics is a reasonable design, but only 15 clinics were included, 57 AI and 49 non-AI clinicians contributed, and the arms show baseline imbalance in clinic region (42.8% vs 34.2% in Thika Road Corridor). The GEE and modified Poisson models adjust for clinic and patient covariates, but they cannot adjust for unmeasured clinician-level differences in documentation habits or baseline error rates. I recommend reporting a baseline or pre-period analysis (e.g., from shadow-mode data or from a pre-study chart sample) to show that the two groups had comparable documentation quality before the intervention, or explicitly acknowledging that the reported effect could partly reflect clinician-level confounding.
  3. [Section 4.1, Fig. 6b, Discussion] The claimed learning effect is circular. The "started red" rate is computed from AI Consult's own red/yellow/green classifications, which are generated by the same model whose alerts are the intervention. A decline in this rate among AI-group clinicians could reflect changes in documentation style that the model no longer flags, rather than genuine improvement in clinical reasoning. The non-AI group's "started red" rate is computed from shadow-mode calls, but prompt and threshold changes during the induction period (documented in Section 3.3) make the metric unstable over time. Please corroborate the learning claim with physician-rated error rates stratified by study week or with a pre-specified test of whether the decline in started-red rate tracks an independent clinical outcome.
  4. [Section 4.1, Table 3] The extrapolation to "22,000 fewer diagnostic errors and 29,000 fewer treatment errors annually at Penda" is not supported by the measurement. The outcomes are physician-rated documentation errors, not confirmed clinical errors, and the absolute numbers apply the relative risk reduction to all 400,000 annual visits without accounting for the fact that the primary endpoint is a rating construct with fair inter-rater reliability. The NNTs likewise quantify the number of visits needed to avoid one rated documentation error, not one patient harm event. Please reframe these projections as "rated errors" and add a caveat that they inherit the uncertainty of the rating-based outcome.
minor comments (4)
  1. [Figure 26] The text in Section 4.2 refers to Fig. 26 as showing clinical note length over time, but the caption in Appendix D.7 describes it as showing the rate of clinician thumbs-up feedback on AI Consult responses. Please correct the figure number or caption.
  2. [Section 3.8, footnote 5] The handling of missing structured chief-complaint data after April 9 is described in a footnote, but it is not clear whether the history RRR in Table 3 is restricted to visits before April 10 or how many visits were excluded. Please state this explicitly in the main text or in the Table 3 notes.
  3. [Section 4.2] The term "attending time" is used without an explicit definition; please clarify whether it is the total visit duration, clinician idle time, or another measure, and state how it is recorded in the EMR.
  4. [References] The KLAS reference lists the year as 2003 while the URL and text suggest a 2023 report; please verify and correct the citation year.

Circularity Check

2 steps flagged · score 2.0 of 10

Primary endpoint uses independent physician raters, but two secondary analyses are self-referential: the learning effect is measured by AI Consult's own red flags, and OpenAI LLMs rate documentation coached by an OpenAI LLM.

  1. other [Section 4.1, 'Clinicians in the AI group learned to avoid common mistakes over time' (Fig. 6b)]
    "We also examine the proportion of visits where AI Consult started red–that is, where the first AI call for any category was red. In the AI group, this rate drops from 45% at the start of the study to 35% at the end of the study, while staying steady at 45-50% in the non-AI group during the study (Fig. 6b). This suggests that AI Consult is training clinicians to avoid common mistakes even prior to AI Consult alerts."

    The 'learning effect' claim uses AI Consult's own red-flag classifications as the outcome metric. The intervention and the measurement are produced by the same model: a declining started-red rate means the model's alerts changed, not independently that clinical errors changed. Without an external anchor to patient outcomes or physician judgments, the decrease in the AI group could reflect model calibration drift, prompt iteration during the induction period, or clinicians learning to satisfy the model's documentation preferences rather than improving care. This is a secondary analysis and does not drive the main physician-rated endpoint.

  2. other [Section 3.8 'LLM rater analysis' and Section 5.1 Limitations]
    "While greater effect sizes may be the result of Goodhart's Law (clinician documentation is assessed by an LLM in AI Consult as well), the greater model-physician agreement compared to physician-physician agreement suggests that LLM ratings, if validated via physician agreement on a subset of cases, may be a way to scale up both routine quality improvement and studies like this one."

    The LLM-rater analysis presents GPT-4.1 and o3 effect sizes as robustness for the main finding, but these raters are OpenAI LLMs grading documentation that a different OpenAI LLM (GPT-4o in AI Consult) actively coached. The paper itself concedes the larger effect sizes may reflect Goodhart's law: the documentation was shaped to satisfy an LLM grader, so an LLM grader is not an independent adjudicator. This analysis is secondary; the primary claim rests on the 108-physician panel.

full rationale

The paper's central comparison (16% fewer diagnostic errors and 13% fewer treatment errors) rests on ratings by 108 independent physicians blinded to group assignment, with golden examples, training, and dual rating of about 25% of visits. That endpoint is not AI Consult's own output, so the main claim is not circular by construction. The skeptical concern that AI Consult coaches documentation completeness—and that Likert 1/2 error definitions include missing history details, missing investigations, and missing additional diagnoses—is a genuine measurement-validity threat, but it is a construct-validity concern rather than a logical reduction of the prediction to the model's inputs; I therefore do not count it as formal circularity. Two secondary analyses are self-referential: the 'learning effect' uses AI Consult's own started-red rate as the outcome, and the LLM-rater analysis uses OpenAI LLMs to grade documentation that an OpenAI LLM helped produce, with the paper itself noting Goodhart's law as a possible explanation for the larger effect sizes. These are exploratory and do not drive the primary endpoint, so the overall circularity score is 2.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the chart-review measurement assumption and the randomization assumption rather than on fitted parameters. One tuning parameter (alert thresholds) shapes the intervention, and four domain assumptions are required.

free parameters (1)
  • Red/yellow/green severity thresholds = not reported (tuned)
    Thresholds for what triggers a red, yellow, or green alert were set through live testing and adjusted during the induction period. They determine which alerts clinicians see and thus shape the intervention itself.
assumptions (4)
  • domain assumption Physician chart-review Likert ratings are a valid measure of clinical error.
    The primary endpoint is physician-assigned Likert scores on documentation; this assumes documentation quality tracks real clinical decision quality.
  • domain assumption Clinician-level randomization within clinics produced exchangeable groups.
    Some randomized clinicians left before rollout and were excluded; Table 1 shows regional imbalance (42.8% vs 34.2% in the Thika Road Corridor) that could reflect non-random attrition.
  • domain assumption No substantial contamination between AI and non-AI clinicians in the same clinic.
    AI and non-AI clinicians worked side by side in the same 15 clinics; the paper does not measure or adjust for spillover effects such as peer learning from AI users.
  • standard math GEE and modified Poisson models are correctly specified.
    Exchangeable correlation and log link are assumed; results are similar across models, but the validity depends on standard statistical assumptions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-based Clinical Decision Support for Primary Care: A Real-World Study." pith.science (2026). https://pith.science/paper/5BZKISFN

@misc{pith2026250716947,
  author       = {Pith},
  title        = {Pith review of: AI-based Clinical Decision Support for Primary Care: A Real-World Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5BZKISFN}},
  note         = {Machine review of arXiv:2507.16947}
}
read the original abstract

We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya, we studied AI Consult, a tool that serves as a safety net for clinicians by identifying potential documentation and clinical decision-making errors. AI Consult integrates into clinician workflows, activating only when needed and preserving clinician autonomy. We conducted a quality improvement study, comparing outcomes for 39,849 patient visits performed by clinicians with or without access to AI Consult across 15 clinics. Visits were rated by independent physicians to identify clinical errors. Clinicians with access to AI Consult made relatively fewer errors: 16% fewer diagnostic errors and 13% fewer treatment errors. In absolute terms, the introduction of AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda alone. In a survey of clinicians with AI Consult, all clinicians said that AI Consult improved the quality of care they delivered, with 75% saying the effect was "substantial". These results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake. We hope this study demonstrates the potential for LLM-based clinical decision support tools to reduce errors in real-world settings and provides a practical framework for advancing responsible adoption.

Figures

Figures reproduced from arXiv: 2507.16947 by the authors.

Figure 1
Figure 1. AI Consult is a safety net that runs in the background of a patient visit to identify potential [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Timeline of AI Consult deployment and quality [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Flow diagram showing visit eligibility, consent, group assignment, and data availability. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (24 more)
Figure 4
Figure 4. Figure 4: Clinical error rates for history-taking, investigations, diagnosis, and treatment, comparing the AI [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Rates of selected clinical failure modes in the AI group compared to the non-AI group. Error bars [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: Rates of visits left in red and started in red over time for AI and non-AI groups. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Clinician survey results: impact of AI Consult on quality of care. [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Median clinician attending time and rate of treatment errors by number of AI Consult triggers [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]
Figure 9
Figure 9. Figure 9: Image of AI Consult yellow notification. [PITH_FULL_IMAGE:figures/full_fig_p035_9.png]
Figure 10
Figure 10. Figure 10: Image of AI Consult yellow popup, after clicking on the notification bell. [PITH_FULL_IMAGE:figures/full_fig_p035_10.png]
Figure 11
Figure 11. Figure 11: Image of AI Consult red popup [PITH_FULL_IMAGE:figures/full_fig_p036_11.png]
Figure 12
Figure 12. Figure 12: Image of AI Consult green notification. 36 [PITH_FULL_IMAGE:figures/full_fig_p036_12.png]
Figure 13
Figure 13. Figure 13: Image of AI Consult green popup, after clicking on the notification bell. [PITH_FULL_IMAGE:figures/full_fig_p037_13.png]
Figure 14
Figure 14. Figure 14: Confusion matrices showing the ratings of two independent raters for each Likert question for the [PITH_FULL_IMAGE:figures/full_fig_p046_14.png]
Figure 15
Figure 15. Figure 15: Likert 1 and 2 rates for history-taking, investigations, diagnosis, and treatment: cases with at [PITH_FULL_IMAGE:figures/full_fig_p047_15.png]
Figure 16
Figure 16. Figure 16: Likert 1 and 2 rates for history-taking, investigations, diagnosis, and treatment - results from [PITH_FULL_IMAGE:figures/full_fig_p047_16.png]
Figure 17
Figure 17. Figure 17: Rate of visits where the final call for the vitals and chief complaint or clinical notes bucket is red, [PITH_FULL_IMAGE:figures/full_fig_p058_17.png]
Figure 18
Figure 18. Figure 18: Rate of visits where the first call for the treatment bucket is red, for AI and non-AI groups over [PITH_FULL_IMAGE:figures/full_fig_p058_18.png]
Figure 19
Figure 19. Figure 19: Rate of visits where the final call for the vitals and chief complaint or clinical notes bucket is red, [PITH_FULL_IMAGE:figures/full_fig_p059_19.png]
Figure 20
Figure 20. Figure 20: Rate of visits where the final call for any of the AI Consult buckets is red, for the AI group [PITH_FULL_IMAGE:figures/full_fig_p059_20.png]
Figure 21
Figure 21. Figure 21: Likert 1 and 2 rates for history-taking, investigations, diagnosis, and treatment, comparing the AI [PITH_FULL_IMAGE:figures/full_fig_p060_21.png]
Figure 22
Figure 22. Figure 22: Likert 1 and 2 rates for history-taking, investigations, diagnosis, and treatment, comparing the [PITH_FULL_IMAGE:figures/full_fig_p061_22.png]
Figure 23
Figure 23. Figure 23: AI group satisfaction net promoter score of AI Consult. [PITH_FULL_IMAGE:figures/full_fig_p077_23.png]
Figure 24
Figure 24. Figure 24: AI group satisfaction with AI Consult. 77 [PITH_FULL_IMAGE:figures/full_fig_p077_24.png]
Figure 25
Figure 25. Figure 25: Mean treatment Likert from GPT-4.1 vs total clinician attending time, binned to 5-minute intervals, in the non-AI and AI groups. 95% CIs calculated with 1000 bootstrap samples. Includes only visits with duration 30 minutes or less [PITH_FULL_IMAGE:figures/full_fig_p0…
Figure 26
Figure 26. Figure 26: Rate of clinician thumbs up feedback on AI Consult responses in the AI group over time. [PITH_FULL_IMAGE:figures/full_fig_p078_26.png]
Figure 27
Figure 27. Figure 27: Rate of clinician thumbs up feedback on AI Consult responses in the AI group over time. [PITH_FULL_IMAGE:figures/full_fig_p079_27.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. A global log for medical AI

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.

  2. Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

    cs.AI 2026-07 accept novelty 5.0 of 10

    Passing medical exams does not make LLMs safe for autonomous triage: they fail to seek missing red flags, and current benchmarks do not test that.

Reference graph

Works this paper leans on

48 extracted references · 46 canonical work pages · cited by 2 Pith papers

  1. [1]

    Evaluate the patient’s chief complaint and vital signs (and MUAC for children ages 6 months–5 years)

  2. [2]

    Determine whether there are urgent or concerning findings that may indicate a medical emergency (Red), incomplete or suboptimal documentation or potential concerns (Yellow), or if everything is appropriate and non-urgent (Green)

  3. [3]

    Severity Thresholds

    Provide concise, actionable recommendations to improve patient safety and care quality. Severity Thresholds

  4. [4]

    patient is well-appearing, normal affect

    Actionable Recommendations •If Red: Offer urgent steps (e.g., re-check vitals, immediate advanced care for suspected emergencies). •If Yellow: Suggest needed clarifications or missing vitals. •If Green: Encourage routine next steps; no critical gaps. Output Structure You must return exactly one severity level in JSON, with an explanatory Reason and an Act...

  5. [5]

    Medications are present but inappropriate

    Likely inappropriate class of antibiotics used (e.g., amox/clavulanic acid used when amoxicillin is appropriate) *if you select this choice, you should also select choice “Medications are present but inappropriate”

  6. [6]

    Referrals are present but inappropriate 8

    Referrals are missing 7. Referrals are present but inappropriate 8. Needed procedures are missing 9. Procedures are present but inappropriate 10. Needed escalations of care are missing 11. Escalations of care are present but inappropriate 12. None of the above Additional resources for physicians Common local brand name pharmaceutical list that physicians ...

  7. [7]

    Chief complaint of severe headache with BP 180/110 mmHg—possible hypertensive emergency

    Red •Potential emergency based on the chief complaint and abnormal vitals (e.g., severe chest pain + very high BP, or severe headache + hypertensive crisis). •All vitals are missing (critical omission). •If a pregnant patient’s complaint and vitals suggest a severe complication (e.g., very high BP, severe edema, etc.). •Example: “Chief complaint of severe...

  8. [8]

    •Some essential vitals are missing but not all

    Yellow •Concerning chief complaint (e.g., chest pain) but vitals do not clearly indicate an emergency; additional assessment is needed. •Some essential vitals are missing but not all . •Respiratory complaints without documented SpO2 or Respiratory Rate . •If the MUAC or other vital sign is borderline, or mild abnormalities that need follow-up but are not emergent

Show all 48 references
  1. [9]

    Vitals within normal limits, mild sore throat, no red flags

    Green •All relevant vitals are documented, no signs of emergent danger in the chief complaint or vitals. •Example: “Vitals within normal limits, mild sore throat, no red flags.” Key Principles

  2. [10]

    •Children under 12 : Temperature, Pulse, (BP is not expected), and for ages 6 months–5 years, MUAC is recommended but not mandatory 80 •Pregnant Patients : BP is crucial

    Essential Vitals •Adults : Temperature, Pulse (HR), Blood Pressure (BP), Height, Weight, and Calculated BMI, and, if respiratory complaints, recommend SpO2. •Children under 12 : Temperature, Pulse, (BP is not expected), and for ages 6 months–5 years, MUAC is recommended but no...

  3. [11]

    MUAC Interpretation •Red : Severe malnutrition (urgent) •Yellow : Moderate malnutrition •Green : No malnutrition •Do not request MUAC for ages outside 6 months–5 years unless specifically indicated

  4. [12]

    Respiratory Rate •While helpful, respiratory rate is not critical for all patients (except those with respiratory complaints, in which case missing RR or SpO2 triggers Red)

  5. [14]

    •No critical tests are missing; no irrelevant or unjustified tests are ordered

    Green •The investigations ordered are appropriate and comprehensive for the clinical scenario. •No critical tests are missing; no irrelevant or unjustified tests are ordered. •Example: A strep test ordered for a patient with sore throat and exudative tonsillitis, or a urine di...

  6. [15]

    •OR there is at least one low-value or marginally justified test ordered

    Yellow •Some recommended investigations are missing or questionable based on the history / exam, but not so critical as to seriously endanger the patient. •OR there is at least one low-value or marginally justified test ordered. •Example: Mild pallor noted but no full haemogra...

  7. [16]

    Add test X,

    Red •Essential diagnostic investigations are omitted, posing a risk of delayed or inaccurate diagnosis. •Clearly inappropriate tests are ordered, showing a major mismatch with the documented presentation. •Example: A patient with severe chest pain but no cardiac or respiratory...

  8. [17]

    Evaluate the clinician’s diagnosis against the visit’s documentation (patient history, exam findings, vitals, labs, etc.)

  9. [18]

    Assess if the diagnosis is appropriate, missing, incomplete, or incorrectly severe given the local epidemi- ology and available resources

  10. [19]

    Severity Thresholds

    Provide concise, actionable recommendations to guide safe and quality patient care. Severity Thresholds

  11. [20]

    •No significant mismatch with history, vitals, labs, or local context

    Green •The listed diagnosis (or diagnoses) accurately reflects the clinical documentation. •No significant mismatch with history, vitals, labs, or local context. •The clinician may safely proceed with management of these diagnoses. •Example: If the patient presents with dysuri...

  12. [21]

    – Additional testing or more thorough documentation is advisable before finalizing

    Yellow •The listed diagnosis broadly aligns with the documentation, but: – There is some uncertainty or missing details preventing a definitive conclusion (e.g., possible severe pathol- ogy but not fully confirmed). – Additional testing or more thorough documentation is advisa...

  13. [22]

    simple cystitis

    Red •A serious mismatch: The listed diagnosis is incompatible with the clinical findings, or a critical diagnosis is missing. •Could result in dangerous consequences if not corrected. •Severe diagnoses listed are not supported by the presentation, or a severe condition is clea...

  14. [23]

    She has had these symptoms before and was diagnosed with UTI

    Green Example Age: 25y Gender: Female Vitals: Temperature: 37.80°C Pulse: 80 bpm Blood Pressure: 120/78 Respiratory Rate: 18 SPO2: 99 Chief Complaint: Dysuria, urinary frequency Clinical notes: Pt complains of dysuria and urinary frequency x2 days. She has had these symptoms b...

  15. [24]

    Yellow Example Clinical Documentation: Age: 16y Gender: Male Vitals: Temperature: 38.50°C Pulse: 90 bpm Blood Pressure: 110/70 Respiratory Rate: 20 SPO2: 98 Chief Complaint: Right lower quadrant abdominal pain, mild nausea Physical Exam: Mild tenderness in RLQ but no rebound o...

  16. [25]

    Red Example 90 Clinical Documentation Age: 35y Gender: Female Vitals: Temperature: 39.20°C Pulse: 105 bpm Blood Pressure: 130/85 Respiratory Rate: 22 SPO2: 98 Chief Complaint: Flank pain, fever, nausea Physical Exam: Notable costovertebral angle tenderness Lab Results: WBC cou...

  17. [26]

    Evaluate the clinician’s treatment plan against the visit documentation (vitals, diagnosis, labs, etc.)

  18. [27]

    Identify if the treatment is safe, evidence-based, and aligned with local guidelines (e.g., MoH Kenya, IMNCI/WHO)

  19. [28]

    Severity Thresholds

    Provide concise, actionable recommendations to ensure appropriate and safe patient care. Severity Thresholds

  20. [29]

    Red •Serious mismatch between treatment and diagnosis. •Unsafe or unnecessary medications (e.g., antibiotics for a confirmed viral illness, sedating antihistamines in young children, monteleukast for respiratory infections without asthma). •Omission of essential medications wh...

  21. [30]

    •Some prescriptions listed are of dubious value to the patient (e.g., cough syrups)

    Yellow •Treatment plan mostly aligns with the documented diagnosis, but: •Minor adjustments to dosage/duration are recommended, or while the medication choice is acceptable, it is not considered a first-line treatment for the condition. •Some prescriptions listed are of dubiou...

  22. [31]

    •No critical omissions or unnecessary interventions

    Green •Treatment plan is complete, accurate, and in compliance with relevant guidelines. •No critical omissions or unnecessary interventions. Specific Guidelines to note

  23. [32]

    Key IMNCI/WHO Guidance for Dehydration in Children ¡5 Years

  24. [33]

    •Then 70 mL/kg over 2.5 hours (¿ 12 months) or 5 hours (¡ 12 months)

    Severe Dehydration •IV Ringer’s Lactate at 30 mL/kg over 30 min (if child ¿ 12 months) or 60 min (¡ 12 months). •Then 70 mL/kg over 2.5 hours (¿ 12 months) or 5 hours (¡ 12 months). •If IV access is not possible, ORS via nasogastric tube at 120 mL/kg over 6 hours

  25. [34]

    Some Dehydration •Oral Rehydration Solution (ORS) at 75 mL/kg over 4 hours. 92

  26. [35]

    All cases should receive zinc supplementation

    No Dehydration •ORS 10 mL/kg after each loose stool. All cases should receive zinc supplementation

  27. [36]

    Septrin (cotrimoxazole) is not a recommended first-line treatment due to its use in TB management

    Urinary Tract Infection Management In Kenya, Nitrofurantoin and Cephalosporins are appropriate first-line therapy for management of UTI in adults and pregnant women. Septrin (cotrimoxazole) is not a recommended first-line treatment due to its use in TB management

  28. [37]

    5”, say “so you’re feeling much better

    Note that Zefcolin (brand name) is a cough syrup and not a cephalosporin antibiotic; it can be used to relieve cough symptoms associated with upper respiratory tract infection in adults and children over 2 years. Output Structure You must return exactly one severity level in J...

  29. [38]

    ∗Please make sure to flag severe outcomes like hospital admission, ICU admission, or death here

    If they do, confirm their answer verbally! – Ask this question even if the patient is feeling better!They may be feeling better because they have already gone to another hospital or chemist 95 – If the patient plans to visit another clinic but hasn’t yet, answer “No”!This ques...

  30. [39]

    Key details in the history are missing (e.g., characterization of chief complaint is lacking key elements such as onset, duration, or associated symptoms, etc

    Chief complaint is absent 2. Key details in the history are missing (e.g., characterization of chief complaint is lacking key elements such as onset, duration, or associated symptoms, etc. or pertinent medical history such as travel/sexual/family history are missing when they ...

  31. [40]

    Documentation of relevant systems on physical exam are absent (e.g., respiratory exam in a patient with cough, description of the rash in a patient with skin findings)

  32. [41]

    Blood group – self request

    Pertinent vital signs are absent 5. None of the above Investigations form and questions Task Specific Clinical Note Investigations: Likert Score : Grade the investigations ordered (or lack of them) on appropriateness, given the clinical documentation. Keep in mind that not all...

  33. [42]

    Unjustified investigations are ordered

  34. [43]

    elevated blood pressure reading

    None of the above (Any investigations ordered are indicated and no key investigations are missing) Diagnosis form and questions Task Specific Clinical Note Diagnosis Likert Score: Grade whether the diagnosis (including primary diagnosis, any additional diagnoses, or the listed...

  35. [44]

    allergic rhinitis

    Primary diagnosis is too specific to be supported based on current documentation or investigations (e.g., using “allergic rhinitis” as the diagnosis rather than “rhinitis”, where it’s clear that rhinitis is present but documentation does not support whether it is a viral, bact...

  36. [45]

    Clinically relevant additional diagnosis is missing (e.g

    Additional diagnosis is likely incorrect 5. Clinically relevant additional diagnosis is missing (e.g. malnutrition) 6. None of the above (All diagnoses are likely correct and no clinically relevant diagnoses are missing) Treatment form and questions Task Specific Clinical Note...

  37. [46]

    Medications are present but inappropriate

    Likely inappropriate use of antibiotics overall (e.g., antibiotics are given for a likely viral infection) *if you select this choice, you should also select choice “Medications are present but inappropriate”

  38. [2006]

    doi: 10.1186/1748-5908-1-1

    ISSN 1748-5908. doi: 10.1186/1748-5908-1-1. 29 S. L. Fleming, A. Lozano, W. J. Haberkorn, J. A. Jindal, E. Reis, R. Thapa, L. Blankemeier, J. Z. Genkins, E. Steinberg, A. Nayak, et al. Medalign: A clinician-generated dataset for instruction following with electronic medical re...

  39. [2018]

    doi: 10.1001/jama.2017.18391. S. Bedi, H. Cui, M. Fuentes, A. Unell, M. Wornow, J. M. Banda, N. Kotecha, T. Keyes, Y. Mai, M. Oez, et al. Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025. M. Benary, X. D. Wang, M. Sc...

  40. [2020]

    elevated blood pressure reading

    URLhttps://www.ajol.info/index.php/eamj/article/view/205283. D. McDuff, M. Schaekermann, T. Tu, and et al. Towards accurate differential diagnosis with large language models.Nature, 626:102–118, 2025. doi: 10.1038/s41586-025-08869-4. B. Middleton, D. F. Sittig, and A. Wright. ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.