Pith. sign in

REVIEW 3 major objections 3 minor 2 cited by

On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper reports evidence that selective prediction, an AI safety mechanism that withholds low-confidence predictions, changes clinician error patterns even when overall accuracy is roughly unchanged.

desk verdict A user study that may falsify a core assumption of selective prediction, but the key error-shift finding needs inferential statistics and a same-case comparison before I'd trust it. read the letter →

arxiv 2508.07617 v1 pith:5FM6MSCR submitted 2025-08-11 cs.HC cs.AI

classification cs.HCcs.AI
keywords selectivepredictionabstentionclinicaldecisionmakinghuman-AIcollaborationautomationbiaserroranalysisunderdiagnosisundertreatment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests a core safety assumption behind selective prediction: when an AI withholds its prediction and tells the user it is abstaining, the user should act as if no AI were present. In a study of 259 clinicians diagnosing and treating simulated hospitalized patients, that assumption fails. Selective prediction recovered most of the accuracy lost to bad AI advice—64% accuracy versus 66% with no AI and 56% with inaccurate AI advice—but it shifted the error profile: clinicians had an 18% increase in missed diagnoses and a 35% increase in missed treatments compared with no AI at all. The authors' case is that abstention is not behaviorally neutral and that accuracy metrics alone hide how AI changes the decisions people make.

What carries the argument

Selective prediction—an AI system that withholds a prediction when its confidence is low and explicitly tells the user it is abstaining—is the intervention under test. The operative comparison is not just overall accuracy but the error mix: the study scores diagnosis and treatment decisions separately across three arms (no AI, inaccurate AI output, and selective prediction with abstention), which is what allows the authors to see that accuracy is nearly recovered while missed diagnoses and treatments rise.

What would settle it

Re-run or audit the study with case-level matching: assign the same vignettes to clinicians with and without AI, measure each vignette's no-AI error rate, and check whether abstention-arm clinicians miss diagnoses and treatments more often on similarly difficult cases. If the abstention versus no-AI difference in missed diagnoses and treatments disappears when difficulty is held fixed, the reported 18% and 35% increases are artifacts of case selection rather than effects of being told the AI abstains.

Watch

Extended reading notes

Core claim

The paper claims that the standard assumption underlying selective prediction—that a user informed of abstention behaves as if no AI were present—is false for this clinical population. Across 259 clinicians, overall decision accuracy under selective prediction (64%) roughly matched the no-AI baseline (66%) and beat the inaccurate-AI condition (56%). But the composition of errors changed: when the AI abstained and said so, clinicians missed more diagnoses (an 18% increase) and more treatments (a 35% increase) than clinicians who saw no AI output at all. The conclusion is that selective prediction can maintain or restore aggregate accuracy while silently worsening omission errors.

Load-bearing premise

The result stands only if the simulated vignettes, the gold-standard diagnosis and treatment labels, and the specific cases on which the AI abstains are representative enough of real deployment that the extra missed diagnoses and treatments in the abstention arm reflect clinician behavior rather than harder cases or an incomplete label set.

Editorial extensions

If this is right

  • Accuracy-only evaluation of abstaining AI is insufficient; safety claims must also track the distribution of missed diagnoses and missed treatments.
  • Telling a user that the AI abstains carries behavioral weight, so selective prediction cannot be treated as a neutral fallback to the no-AI baseline.
  • Clinical deployments of abstaining AI should monitor for omission bias and may need countermeasures such as prompting the clinician to revisit alternatives when abstention is signaled.
  • A 35% increase in missed treatments is a large shift for patient care even when aggregate accuracy looks stable, so deployment dashboards should separate error types.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same abstention-as-silence effect could appear in other high-stakes human-AI settings—radiology, triage, legal or credit review—where a visible abstention might be read as a de facto negative or dismissive signal; this is an extension the paper does not test.
  • If abstention signals are interpreted as 'nothing to see here,' the omission bias could grow as clinicians gain familiarity with a reliable model's abstention behavior; testing that learning curve is a natural next experiment.
  • A design-level remedy not examined here would be to couple abstention with an explicit prompt to consider a differential diagnosis or seek a second opinion, which would directly test whether the error shift is caused by the abstention signal itself.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper tests a core assumption of selective prediction: that when an AI abstains and informs the user, the user's decisions remain the same as they would be without AI involvement. It reports a user study with 259 clinicians randomized across three arms: no AI, inaccurate AI, and selective prediction (AI abstains on unreliable cases and informs the user). The abstract reports overall decision accuracy of 66% (no AI), 56% (inaccurate AI), and 64% (selective prediction), with overlapping 95% confidence intervals. The central claim is that selective prediction nearly preserves overall accuracy but changes the pattern of errors: informed abstention is associated with an 18% increase in missed diagnoses and a 35% increase in missed treatments compared to no AI input. The authors conclude that the behavioral assumption underlying selective prediction is false in this clinical population.

Significance. If the error-shift finding is robust, the paper makes an important contribution: it challenges a standard assumption in selective prediction by showing that informing users about abstention can alter decisions in a way that accuracy metrics obscure. The study is directly relevant to human-AI decision making, clinical decision support, and the safe deployment of abstaining models. The overall accuracy recovery is modest and statistically weak, so the paper's significance rests almost entirely on the secondary error-pattern results. The manuscript would be valuable if it can convincingly support those results with appropriate inference and a clean comparison basis.

major comments (3)
  1. [Abstract] The headline error-shift results — 18% more missed diagnoses and 35% more missed treatments under selective prediction — are reported without confidence intervals, p-values, or effect sizes. Given that the primary accuracy estimates have half-widths of roughly 9–10 percentage points and each arm has about 86 clinicians, these secondary differences may be well within sampling noise. Please report inferential statistics for the missed-diagnosis and missed-treatment rates, state whether these endpoints were pre-specified, and provide the underlying counts or rates by arm.
  2. [Abstract; comparison basis] The claim that informed abstention increases missed diagnoses and treatments is compared against 'no AI input at all.' If the selective-prediction arm's abstained cases are systematically harder or lower-confidence than the average no-AI case, the observed error shift could be an artifact of case mix rather than a behavioral response to abstention. The paper must clarify whether the no-AI comparison is restricted to the same vignette subset on which the AI abstained, or otherwise adjusted for case difficulty. Without this, the central claim is confounded.
  3. [Methods (not verifiable in provided text)] The provided full text is heavily corrupted, so randomization, vignette selection, gold-standard labeling, and the abstention mechanism cannot be verified. Even setting aside the encoding issue, the abstract does not report how the 'inaccurate AI' arm was constructed (error rate, confidence threshold) or whether the selective-prediction abstention threshold was tuned to produce the reported accuracy recovery. These details are necessary to assess whether the selective-prediction condition is a fair representation of the method and whether the comparisons are internally valid.
minor comments (3)
  1. [Abstract] The phrase 'clinician accuracy declined' is based on overlapping confidence intervals (66% vs. 56%); please clarify whether this decline is statistically significant or present it as a point estimate only.
  2. [General] The full text as provided is unreadable due to encoding corruption; the authors should ensure the arXiv source is correctly rendered, as the current version prevents verification of tables, figures, and methods details.
  3. [General] The abstract would benefit from stating the number of cases per clinician and whether outcomes were measured per vignette or per clinician, as this affects the appropriate statistical model for the error-shift comparisons.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper is an empirical user study that tests a behavioral assumption; no derivation reduces to its own inputs.

full rationale

This paper reports a user study in which 259 clinicians were randomly assigned to conditions (no AI, AI without selective prediction, AI with selective prediction) and their diagnostic/treatment accuracy was measured against gold-standard labels. The central claim that informing clinicians of AI abstention changes their error patterns (18% missed-diagnosis increase, 35% missed-treatment increase relative to no AI) is an empirical behavioral finding, not a quantity defined in terms of itself. The authors do not fit a parameter and then 'predict' it; the abstention threshold is a design choice, not a fitted input. No equation equates the outcome with an input. No load-bearing self-citation appears; the selective-prediction assumption is explicitly stated as the hypothesis under test, not as an established theorem. The skeptical concern that abstention cases may be harder than the no-AI full case set, or that the reported increases lack inferential statistics, is a statistical-validity objection, not a circularity. Therefore, under the given rubric, no step reduces to its own inputs, and the score reflects only the absence of circularity, with a minor distinction from 0 due to design limitations that are not circular.

Assumptions & free parameters 2 free parameters · 2 assumptions · 0 invented entities

No new physical or conceptual entities are invented. The paper's dependence is on design choices (AI error distribution, abstention threshold) and domain assumptions (vignette validity, gold-standard label correctness) that cannot be checked from the abstract. The abstention-threshold parameter is the closest analogue to a hand-tuned quantity, and the error-shift result is conditional on it.

free parameters (2)
  • AI error-case selection and accuracy = not reported in abstract
    The 'inaccurate AI predictions' condition requires a chosen set of cases on which the AI errs; its size and composition drive the observed 10-point accuracy drop and the 35% undertreatment shift. This is a design choice the central claim depends on.
  • Selective prediction abstention threshold / coverage = not reported in abstract
    The rate and pattern of abstentions determine how often clinicians are exposed to the abstention cue; the 18% missed-diagnosis and 35% missed-treatment shifts are conditional on this threshold choice.
assumptions (2)
  • domain assumption Clinicians' decisions in the simulated vignettes transfer to real hospital behavior
    The study is framed as a clinical user study and the general lesson is about clinicians in real practice; that generalizes only if the simulated cases elicit realistic decision behavior. Not verifiable from the abstract.
  • domain assumption The gold-standard diagnoses and treatments used for scoring are correct and complete
    Missed diagnoses and missed treatments are computed against these labels; incomplete gold-standard labels would inflate both the 18% and 35% increases and could reverse the error-shift conclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making." pith.science (2026). https://pith.science/paper/5FM6MSCR

@misc{pith2026250807617,
  author       = {Pith},
  title        = {Pith review of: On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5FM6MSCR}},
  note         = {Machine review of arXiv:2508.07617}
}
read the original abstract

AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deployed. These inaccuracies, combined with automation bias, where humans overrely on AI predictions, can result in worse decisions. Selective prediction, in which potentially unreliable model predictions are hidden from users, has been proposed as a solution. This approach assumes that when AI abstains and informs the user so, humans make decisions as they would without AI involvement. To test this assumption, we study the effects of selective prediction on human decisions in a clinical context. We conducted a user study of 259 clinicians tasked with diagnosing and treating hospitalized patients. We compared their baseline performance without any AI involvement to their AI-assisted accuracy with and without selective prediction. Our findings indicate that selective prediction mitigates the negative effects of inaccurate AI in terms of decision accuracy. Compared to no AI assistance, clinician accuracy declined when shown inaccurate AI predictions (66% [95% CI: 56%-75%] vs. 56% [95% CI: 46%-66%]), but recovered under selective prediction (64% [95% CI: 54%-73%]). However, while selective prediction nearly maintains overall accuracy, our results suggest that it alters patterns of mistakes: when informed the AI abstains, clinicians underdiagnose (18% increase in missed diagnoses) and undertreat (35% increase in missed treatments) compared to no AI input at all. Our findings underscore the importance of empirically validating assumptions about how humans engage with AI within human-AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence

    cs.LG 2026-08 conditional novelty 5.0 of 10

    CRS-Triage, which jointly models modality reliability, cross-modal consistency, and a learned confidence score, improves triage accuracy and reduces under-triage on MIMIC-IV-ED compared with evidential and fusion baselines.

  2. SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

    cs.LG 2026-07 conditional novelty 5.0 of 10

    An event-graph GNN plus conformal FDR selection can impute irregular clinical labs and release only a subset with controlled rates of clinically large errors.

Reference graph

Works this paper leans on

50 extracted references · 41 canonical work pages · cited by 2 Pith papers

  1. [1]

    E.; Sridharan, A.; Soleimani, H.; Zhan, A.; Rawat, N.; Johnson, L.; Hager, D

    Adams, R.; Henry, K. E.; Sridharan, A.; Soleimani, H.; Zhan, A.; Rawat, N.; Johnson, L.; Hager, D. N.; Cosgrove, S. E.; Markowski, A.; et al. 2022. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature medicine, 28(7): 1455--1460

  2. [2]

    F.; and Ribeiro, L

    Ara \'u jo, B.; Gomes, S. F.; and Ribeiro, L. 2024. Critical thinking pedagogical practices in medical education: a systematic review. Frontiers in Medicine, 11: 1358444

  3. [3]

    Banovic, N.; Yang, Z.; Ramesh, A.; and Liu, A. 2023. Being trustworthy is not enough: How untrustworthy artificial intelligence (AI) can deceive the end-users and gain their trust. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1): 1--17

  4. [4]

    T.; and Weld, D

    Bansal, G.; Wu, T.; Zhou, J.; Fok, R.; Nushi, B.; Kamar, E.; Ribeiro, M. T.; and Weld, D. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems, 1--16

  5. [5]

    Beery, S.; Morris, D.; and Yang, S. 2019. Efficient pipeline for camera trap image review. arXiv preprint arXiv:1907.06772

  6. [6]

    Bondi, E.; Koster, R.; Sheahan, H.; Chadwick, M.; Bachrach, Y.; Cemgil, T.; Paquet, U.; and Dvijotham, K. 2022. Role of human-AI interaction in selective prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 5286--5294

  7. [7]

    B.; and Gajos, K

    Bu c inca, Z.; Malaya, M. B.; and Gajos, K. Z. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-computer Interaction, 5(CSCW1): 1--21

  8. [8]

    Chen, J.; Yoon, J.; Ebrahimi, S.; Arik, S.; Jha, S.; and Pfister, T. 2023. ASPEST: Bridging the Gap Between Active Learning and Selective Prediction. arXiv preprint arXiv:2304.03870

Show all 50 references
  1. [9]

    Chow, C. 1970. On optimum recognition error and reject tradeoff. IEEE Transactions on information theory, 16(1): 41--46

  2. [10]

    Clayton, D. G. 1996. Generalized linear mixed models. Markov chain Monte Carlo in practice, 1: 275--302

  3. [11]

    Davison, A. 1997. Bootstrap methods and their application. Cambridge University Press

  4. [12]

    De-Arteaga, M.; Fogliato, R.; and Chouldechova, A. 2020. A case for humans-in-the-loop: Decisions in the presence of erroneous algorithmic scores. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1--12

  5. [13]

    Dong, X.; and Hayes, C. C. 2012. Uncertainty visualizations: Helping decision makers become more aware of uncertainty and its implications. Journal of Cognitive Engineering and Decision Making, 6(1): 30--56

  6. [14]

    El-Yaniv, R.; and Wiener, Y. 2010. On the Foundations of Noise-free Selective Classification. Journal of Machine Learning Research, 11(5)

  7. [15]

    R.; and Sjoding, M

    Farzaneh, N.; Ansari, S.; Lee, E.; Ward, K. R.; and Sjoding, M. W. 2023. Collaborative strategies for deploying artificial intelligence to complement physician diagnoses of acute respiratory distress syndrome. NPJ Digital Medicine, 6(1): 62

  8. [16]

    Z.; and Mamykina, L

    Gajos, K. Z.; and Mamykina, L. 2022. Do people engage cognitively with AI? Impact of AI assistance on incidental learning. In Proceedings of the 27th International Conference on Intelligent User Interfaces, 794--806

  9. [17]

    J.; Lermer, E.; Coughlin, J

    Gaube, S.; Suresh, H.; Raue, M.; Merritt, A.; Berkowitz, S. J.; Lermer, E.; Coughlin, J. F.; Guttag, J. V.; Colak, E.; and Ghassemi, M. 2021. Do as AI say: susceptibility in deployment of clinical decision-aids. NPJ digital medicine, 4(1): 31

  10. [18]

    Gkatzia, D.; Lemon, O.; and Rieser, V. 2016. Natural language generation enhances human decision-making with uncertain information. arXiv preprint arXiv:1606.03254

  11. [19]

    L.; Amaral, L

    Goldberger, A. L.; Amaral, L. A.; Glass, L.; Hausdorff, J. M.; Ivanov, P. C.; Mark, R. G.; Mietus, J. E.; Moody, G. B.; Peng, C.-K.; and Stanley, H. E. 2000. PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals. circul...

  12. [20]

    P.; Chetty, U.; O'Donnell, P.; Gajria, C.; and Blackadder-Weinstein, J

    Gopal, D. P.; Chetty, U.; O'Donnell, P.; Gajria, C.; and Blackadder-Weinstein, J. 2021. Implicit bias in healthcare: clinical practice, research and decision making. Future healthcare journal, 8(1): 40--48

  13. [21]

    C.; Wu, D.; Narayanaswamy, A.; Venugopalan, S.; Widner, K.; Madams, T.; Cuadros, J.; et al

    Gulshan, V.; Peng, L.; Coram, M.; Stumpe, M. C.; Wu, D.; Narayanaswamy, A.; Venugopalan, S.; Widner, K.; Madams, T.; Cuadros, J.; et al. 2016. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. jama, 316...

  14. [22]

    Irvin, J.; Rajpurkar, P.; Ko, M.; Yu, Y.; Ciurea-Ilcus, S.; Chute, C.; Marklund, H.; Haghgoo, B.; Ball, R.; Shpanskaya, K.; et al. 2019. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificia...

  15. [23]

    Jabbour, S.; Fouhey, D.; Kazerooni, E.; Wiens, J.; and Sjoding, M. W. 2022. Combining chest X-rays and electronic health record (EHR) data using machine learning to diagnose acute respiratory failure. Journal of the American Medical Informatics Association, 29(6): 1060--1068

  16. [24]

    S.; Kazerooni, E

    Jabbour, S.; Fouhey, D.; Shepard, S.; Valley, T. S.; Kazerooni, E. A.; Banovic, N.; Wiens, J.; and Sjoding, M. W. 2023. Measuring the impact of AI in the diagnosis of hospitalized patients: a randomized clinical vignette survey study. Jama, 330(23): 2275--2284

  17. [25]

    Johnson, A.; Pollard, T.; Mark, R.; Berkowitz, S.; and Horng, S. 2024. Mimic-cxr database. PhysioNet10, 13026: C2JT1Q

  18. [26]

    E.; Pollard, T

    Johnson, A. E.; Pollard, T. J.; Berkowitz, S. J.; Greenbaum, N. R.; Lungren, M. P.; Deng, C.-y.; Mark, R. G.; and Horng, S. 2019. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, 6(1): 317

  19. [27]

    A.; Abril, M

    Kempker, J. A.; Abril, M. K.; Chen, Y.; Kramer, M. R.; Waller, L. A.; and Martin, G. S. 2020. The epidemiology of respiratory failure in the United States 2002--2017: a serial cross-sectional study. Critical Care Explorations, 2(6): e0128

  20. [28]

    I'm Not Sure, But

    Kim, S. S.; Liao, Q. V.; Vorvoreanu, M.; Ballard, S.; and Vaughan, J. W. 2024. " I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, 822--835

  21. [29]

    Why is' Chicago'deceptive?

    Lai, V.; Liu, H.; and Tan, C. 2020. " Why is' Chicago'deceptive?" Towards Building Model-Driven Tutorials for Humans. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1--13

  22. [30]

    Lai, V.; and Tan, C. 2019. On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In Proceedings of the conference on fairness, accountability, and transparency, 29--38

  23. [31]

    How do I fool you?

    Lakkaraju, H.; and Bastani, O. 2020. " How do I fool you?" Manipulating User Trust via Misleading Black Box Explanations. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 79--85

  24. [32]

    A.; O'Reilly, G.; Kelly, B

    Lambe, K. A.; O'Reilly, G.; Kelly, B. D.; and Curristan, S. 2016. Dual-process cognitive interventions to enhance diagnostic reasoning: a systematic review. BMJ quality & safety, 25(10): 808--820

  25. [33]

    Li, Z.; and Passonneau, R. J. 2024. Joint Training for Selective Prediction. arXiv preprint arXiv:2410.24029

  26. [34]

    Lykouris, T.; and Weng, W. 2024. Learning to defer in content moderation: The human-ai interplay. arXiv preprint arXiv:2402.12237

  27. [35]

    M.; Sieniek, M.; Godbole, V.; Godwin, J.; Antropova, N.; Ashrafian, H.; Back, T.; Chesus, M.; Corrado, G

    McKinney, S. M.; Sieniek, M.; Godbole, V.; Godwin, J.; Antropova, N.; Ashrafian, H.; Back, T.; Chesus, M.; Corrado, G. S.; Darzi, A.; et al. 2020. International evaluation of an AI system for breast cancer screening. Nature, 577(7788): 89--94

  28. [36]

    Mozannar, H.; Lang, H.; Wei, D.; Sattigeri, P.; Das, S.; and Sontag, D. 2023. Who should predict? exact algorithms for learning to defer to humans. In International conference on artificial intelligence and statistics, 10520--10545. PMLR

  29. [37]

    Y.; Kim, I

    Nam, B.; Kim, J. Y.; Kim, I. Y.; and Cho, B. H. 2022. Selective prediction with long short-term memory using unit-wise batch standardization for time series health data sets: algorithm development and validation. JMIR Medical Informatics, 10(3): e30587

  30. [38]

    M.; Carignan, D.; and Horvitz, E

    Nori, H.; King, N.; McKinney, S. M.; Carignan, D.; and Horvitz, E. 2023. Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375

  31. [39]

    R.; Monteiro, S

    Norman, G. R.; Monteiro, S. D.; Sherbino, J.; Ilgen, J. S.; Schmidt, H. G.; and Mamede, S. 2017. The causes of errors in clinical reasoning: cognitive biases, knowledge deficits, and dual process thinking. Academic Medicine, 92(1): 23--30

  32. [40]

    V.; and Banovic, N

    Prabhudesai, S.; Yang, L.; Asthana, S.; Huan, X.; Liao, Q. V.; and Banovic, N. 2023. Understanding Uncertainty: How Lay Decision-makers Perceive and Interpret Uncertainty in Human-AI Decision Making. In Proceedings of the 28th International Conference on Intelligent User Inter...

  33. [41]

    Schemmer, M.; Kuehl, N.; Benz, C.; Bartos, A.; and Satzger, G. 2023. Appropriate reliance on AI advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces, 410--422

  34. [42]

    K.; Das, S.; Panda, R.; Sattigeri, P.; and Wornell, G

    Shah, A.; Bu, Y.; Lee, J. K.; Das, S.; Panda, R.; Sattigeri, P.; and Wornell, G. W. 2022. Selective regression under fairness criteria. In International Conference on Machine Learning, 19598--19615. PMLR

  35. [43]

    A.; Levin, J.; Kahn, J

    Sivaraman, V.; Bukowski, L. A.; Levin, J.; Kahn, J. M.; and Perer, A. 2023. Ignore, trust, or negotiate: understanding clinician acceptance of AI-based treatment recommendations in health care. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 1--18

  36. [44]

    Smith, T. M. 2023. What’s the difference between physician assistants and physicians?

  37. [45]

    Strong, J.; Men, Q.; and Noble, J. A. 2025. Trustworthy and Practical AI for Healthcare: A Guided Deferral System with Large Language Models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 28413--28421

  38. [46]

    Wang, D.-Y.; Ding, J.; Sun, A.-L.; Liu, S.-G.; Jiang, D.; Li, N.; and Yu, J.-K. 2023. Artificial intelligence suppression as a strategy to mitigate artificial intelligence automation bias. Journal of the American Medical Informatics Association, 30(10): 1684--1692

  39. [47]

    V.; and Bellamy, R

    Zhang, Y.; Liao, Q. V.; and Bellamy, R. K. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 conference on fairness, accountability, and transparency, 295--305

  40. [48]

    Zwaan, L.; Thijs, A.; Wagner, C.; van der Wal, G.; and Timmermans, D. R. 2012. Relating faults in diagnostic reasoning with diagnostic errors and patient harm. Academic Medicine, 87(2): 149--156

  41. [49]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  42. [50]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.