Pith. sign in

REVIEW 4 major objections 4 minor 60 references

2-Factor Retrieval for Improved Human-AI Decision Making in Radiology

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that a 2-factor retrieval interface—an AI diagnosis plus four physician-confirmed reference images—raises clinician accuracy on chest X-rays above saliency maps, AI-only, and no-AI baselines, reaching roughly 70 percent…

desk verdict The 2FR idea is simple and worth testing, but the paper's central claim outruns its statistics: no modality-level test is reported and the confidence intervals overlap. read the letter →

arxiv 2412.00372 v1 pith:LUFBIGVU submitted 2024-11-30 cs.HC cs.AI

classification cs.HCcs.AI
keywords 2-factorretrievalverification-basedAIhuman-AIdecisionmakingchestX-raydiagnosisexplainablesaliencymapsclinicianconfidencehuman-computerinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that how an AI prediction is presented to clinicians changes how much it helps. It introduces '2-factor retrieval' (2FR): alongside the AI-predicted diagnosis, the clinician sees four physician-confirmed chest X-ray examples of that same diagnosis. In a study with 69 physicians, 2FR produced the highest diagnostic accuracy of any tested mode, about 70 percent when the AI was correct, outperforming saliency maps, AI-only, and no-AI conditions. The authors interpret this as evidence that a verification-based interface helps clinicians recall and compare the features of a pathology, rather than merely trusting the AI's answer. If true, the result would argue for retrieval-based design in clinical decision support systems.

What carries the argument

The central object is '2-factor retrieval' (2FR), an interface plus retrieval scheme that, given an AI-predicted diagnosis, displays four physician-confirmed reference images of that diagnosis without processing them through the model. It is called 2-factor because the correct images must be retrieved by the AI and the clinician must associate those images with the pathology in the current X-ray, forming a verification loop. The comparison conditions are saliency maps from the same chest X-ray model, AI-only prediction, and no AI; participants chose one of 14 diagnoses and rated confidence on a 10-point scale. Linear mixed-effects models with a random participant intercept estimate how modality, AI correctness, difficulty, specialty, and experience affect accuracy and confidence.

What would settle it

Run the same 12-case study but replace the four physician-confirmed reference images with randomly selected same-label images; if 2FR's accuracy advantage over saliency disappears, the benefit is driven by the specific canonical examples rather than by the verification process itself.

Watch

Extended reading notes

Core claim

On its own terms, this paper discovers that presenting an AI diagnosis alongside four physician-confirmed reference images of that diagnosis yields the highest clinician accuracy among the tested modes of AI assistance. Across 69 physicians reading 12 chest X-rays, 2FR reached about 70 percent accuracy when the AI prediction was correct, compared with roughly 65 percent for saliency maps, 64 percent for AI-only, and 45 percent for no AI. The advantage was strongest for radiologists and for clinicians with less than 11 years of practice, and 2FR was the best-performing modality at low confidence, roughly tripling saliency's accuracy in that subgroup. When the AI was wrong, all modalities fell to about 25 percent, similar to no AI, so the benefit is tied to the AI being right. The authors conclude that a verification-based presentation can improve human-AI diagnostic performance without changing the AI model.

Load-bearing premise

The comparison's validity depends on the four reference images shown for each AI prediction being correctly labeled and genuinely typical of the predicted pathology, yet the paper does not describe the retrieval algorithm or the verification of these images.

Editorial extensions

If this is right

  • Clinical decision support systems could raise diagnostic accuracy by changing only the interface, not the AI model, if 2FR's effect replicates.
  • Low-confidence clinicians, who otherwise struggle most, are the clearest beneficiaries of 2FR-based presentation.
  • 2FR does not compensate for an incorrect AI; accuracy under 2FR falls to no-AI levels when the AI errs, so the method inherits the AI's reliability.
  • Task experts such as radiologists gain more from reference images than from saliency maps, while saliency maps were less helpful for experts than for non-experts.
  • Because confidence stays roughly flat while accuracy moves, self-assessed confidence is not a reliable proxy for whether AI assistance helped.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A straightforward test left implicit in the paper is to vary the quality of the retrieved set (random same-label images versus canonical examples); this would separate the verification mechanism from simple anchoring on any same-label exemplar.
  • If the effect is driven by canonical examples, 2FR could be combined with model confidence or uncertainty estimates so that reference images are only shown when the AI prediction is likely to be correct.
  • The same retrieval-based interface could be tried in other visual diagnostic domains, such as dermatology or ophthalmology, although the paper only reports chest X-rays.
  • The paper's decoupling of confidence from accuracy suggests that future studies should track decision time and verification behavior, not just final diagnosis, as a predictor of when reference images help.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a new human-AI decision support modality called 2-factor retrieval (2FR), in which an AI prediction is presented together with four reference images purportedly showing the same diagnosis, and compares it with saliency-map explanations, AI-only assistance, and no assistance in a chest X-ray diagnosis task. In an online study with 69 physicians, each completing 12 cases (three per modality), the authors report that 2FR yields the highest overall accuracy (about 70% when the AI is correct) and claim particular benefits for radiologists and for low-confidence decisions. The paper argues that 2FR enables a verification-based reasoning step that improves human-AI decision making.

Significance. If the central claim were solidly established, the 2FR idea would be a valuable, low-cost contribution to the human-AI decision-making literature: it requires no new model training and could be paired with any classifier. The study also addresses a relevant gap by evaluating verification-based interfaces against explainability methods in a medical domain. However, the current evidence base is too weak to support the headline accuracy claim: no statistical test of the modality effect is reported, the key confidence intervals overlap broadly, and the subgroup analyses are severely underpowered. The idea is promising, but the paper as written does not meet the evidentiary standard needed to recommend 2FR over existing interfaces.

major comments (4)
  1. [Section 4.1 and Fig. 4] The central claim that 2FR increases clinician accuracy over the other modalities is not supported by any reported statistical test of the modality effect. Section 3.5 states that mixed-effects models for accuracy were constructed, but Section 4.1 reports only the effect of AI correctness (p<0.001) and no pairwise comparisons among 2FR, Saliency, AI-only, and No AI. In Fig. 4, the 95% confidence intervals for 2FR (0.57–0.81), Saliency (0.52–0.77), and AI-only (0.51–0.76) when the AI is correct overlap substantially, and with N=69 and only three trials per modality per participant, the observed point differences (0.69 vs. 0.65 vs. 0.64) are within sampling noise. Without a reported test, the manuscript's headline claim is not statistically verified.
  2. [Section 3.4] The 2FR intervention is underspecified. The text states that four images 'recognized by other physicians to represent that diagnosis' were retrieved, but it does not describe the retrieval algorithm, the source pool from which the images were drawn, or any verification procedure for ensuring that the retrieved images are correctly labeled and representative. Because the mechanism of 2FR depends on the retrieved images being valid examples of the AI-predicted pathology, failing to specify this makes the comparison difficult to interpret and impossible to reproduce. The authors should describe how the reference images were selected and what quality checks were applied to them.
  3. [Section 4.1 and Figs. 5–7] The subgroup claims (e.g., '2FR significantly aids expert users,' '2FR is most useful for clinicians with less experience,' and the low-confidence advantage shown in Fig. 10) are not backed by any reported statistical tests for interaction or subgroup effects. With 25 radiologists and 44 non-radiologists, and only three trials per modality per participant, the cell sizes for the experience-by-modality and confidence-by-modality analyses are very small; the differences highlighted in the text (e.g., the 3x/2x low-confidence comparison) could be driven by a small number of responses. The manuscript should report interaction tests or explicitly label these analyses as exploratory.
  4. [Section 4.1] The interpretation that 'physicians are overly trusting of AI predictions' is confounded by case difficulty. The authors acknowledge that cases with incorrect AI predictions might be inherently harder, but the conclusion that clinicians over-trust AI is drawn from observing that accuracy tracks AI correctness. Because AI correctness is a property of the case, not of the AI assistance, and is deliberately set at 66.7%, this correlation alone does not demonstrate over-reliance. The alignment-based measure described in Section 3.7 (the correlation between physician and AI diagnoses) would be a more direct test, but it is not reported.
minor comments (4)
  1. [Section 3.4] The 14 diagnosis options given to participants are not listed; the full set of choices should be provided so readers can assess the difficulty and granularity of the task.
  2. [Section 3.4] The 'Easy' versus 'Hard' case labels are not defined; the manuscript does not state who assigned these ratings or what criteria were used, which is important because the paper interprets accuracy differences across these categories.
  3. [Section 4.2] The claim that clinician confidence is 'marginal' across conditions is stated without reporting the mixed-effects model results for confidence that are promised in Section 3.5; the authors should include the relevant statistical output or state that confidence analyses were descriptive only.
  4. [Fig. 3 and Fig. 4] Fig. 3 would benefit from error bars, and Fig. 4 is presented as a table rather than a figure; the manuscript should be consistent in labeling and should state explicitly that Fig. 4 reports 95% confidence intervals.

Circularity Check

0 steps flagged · score 0.0 of 10

No material circularity: this is an empirical human-subjects comparison, not a derivation chain.

full rationale

The paper reports a between-subjects experiment comparing four decision-support modalities: AI-only, saliency, 2FR, and no AI. The central claim is that 2FR increases clinician accuracy. This is an empirical claim about an interface intervention, not a derived or fitted prediction. The 2FR condition retrieves four images using the AI's predicted label, but the claim is not that the retrieved images verify the AI's prediction; rather, it is that displaying such images changes clinician behavior. No parameter is fitted to a subset of the data and then renamed as a prediction, and no equation is shown to be equivalent to an input by construction. The paper does cite prior work for the NIH ChestX-ray dataset and its saliency maps [49], but that is external data/benchmark evidence, not self-citation. The absence of a reported modality-level significance test and the overlapping confidence intervals in Fig. 4 are evidentiary/statistical-support concerns, not circularity. The paper even acknowledges a possible confound (cases with incorrect AI predictions may be inherently harder), which further shows the authors treat the comparison as an empirical question rather than a definitional one. Therefore, no circular step is exhibited, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the correctness of the dataset labels, the representativeness of the 2FR reference images, the saliency baseline, and the validity of the confidence measure. No equations or fitted parameters are involved.

assumptions (4)
  • domain assumption The ChestX-ray8 dataset labels used as ground truth are accurate.
    Section 3.4 uses 24 images from the NIH ChestX-ray dataset, with the dataset's pathology labels treated as the correct diagnosis for scoring physician accuracy.
  • domain assumption The four reference images shown in the 2FR condition are correctly labeled and representative of the AI-predicted diagnosis.
    Section 3.4 states 2FR shows 'four more images recognized by other physicians to represent that diagnosis' but provides no retrieval or verification details.
  • domain assumption The DCNN from [49] and its saliency maps are a representative explainable AI baseline.
    Section 3.4 uses the saliency maps from [49] as the benchmark explainability technique.
  • domain assumption Self-reported confidence on a 10-point Likert scale is a valid measure of decision confidence.
    Section 3.7 measures confidence with a single 10-point item, which is treated as a reliable dependent variable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of 2-Factor Retrieval for Improved Human-AI Decision Making in Radiology." pith.science (2026). https://pith.science/paper/LUFBIGVU

@misc{pith2026241200372,
  author       = {Pith},
  title        = {Pith review of: 2-Factor Retrieval for Improved Human-AI Decision Making in Radiology},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LUFBIGVU}},
  note         = {Machine review of arXiv:2412.00372}
}
read the original abstract

Human-machine teaming in medical AI requires us to understand to what degree a trained clinician should weigh AI predictions. While previous work has shown the potential of AI assistance at improving clinical predictions, existing clinical decision support systems either provide no explainability of their predictions or use techniques like saliency and Shapley values, which do not allow for physician-based verification. To address this gap, this study compares previously used explainable AI techniques with a newly proposed technique termed '2-factor retrieval (2FR)', which is a combination of interface design and search retrieval that returns similarly labeled data without processing this data. This results in a 2-factor security blanket where: (a) correct images need to be retrieved by the AI; and (b) humans should associate the retrieved images with the current pathology under test. We find that when tested on chest X-ray diagnoses, 2FR leads to increases in clinician accuracy, with particular improvements when clinicians are radiologists and have low confidence in their decision. Our results highlight the importance of understanding how different modes of human-AI decision making may impact clinician accuracy in clinical decision support systems.

Figures

Figures reproduced from arXiv: 2412.00372 by the authors.

Figure 1
Figure 1. Different Modes of AI-Human Decision Making, including [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Example interface shown to radiologists. Panel A demonstrates 2FR, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Physician accuracy across AI correctness. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Standard error and confidence intervals of physician accuracy across [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Accuracy across modes of AI-Human decision making and AI cor [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Accuracy across modalities and AI correctness based on clinician [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 8
Figure 8. Figure 8: Clinician confidence across modalities and AI correctness. [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: Confidence across modalities and AI correctness based on chest x-ray [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 46 canonical work pages

  1. [1]

    Lamia Alam and Shane Mueller. 2021. Examining the effect of explanation on satisfaction and trust in AI diagnostic systems. BMC medical informatics and decision making 21, 1 (2021), 178

  2. [2]

    Catarina Barata, Veronica Rotemberg, Noel CF Codella, Philipp Tschandl, Christoph Rinner, Bengu Nisa Akay, Zoe Apalla, Giuseppe Argenziano, Allan Halpern, Aimilios Lallas, et al. 2023. A reinforcement learning model for AI-based decision support in skin cancer. Nature Medicine 29, 8 (2023), 1941–1946

  3. [3]

    Mustafa Bilgic and Raymond J Mooney. 2005. Explaining recommendations: Satisfaction vs. promotion. In Beyond personalization workshop, IUI , Vol. 5. 153

  4. [4]

    Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems. In Proceedings of the 25th international conference on intelligent user interfaces. 454–464

  5. [5]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (2021), 1–21

  6. [6]

    Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The role of explanations on trust and reliance in clinical decision support systems. In 2015 international conference on healthcare informatics . IEEE, 160–169

  7. [7]

    Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry

  8. [8]

    Srikant Devaraj, Sushil K Sharma, Dyan J Fausto, Sara Viernes, Hadi Kharrazi, et al. 2014. Barriers and facilitators to clinical decision support systems adoption: a systematic review. Journal of Business Administration Research 3, 2 (2014), 36

Show all 60 references
  1. [9]

    Upol Ehsan, Philipp Wintersberger, Q Vera Liao, Martina Mara, Marc Streit, Sandra Wachter, Andreas Riener, and Mark O Riedl. 2021. Operationalizing human- centered perspectives in explainable AI. In Extended abstracts of the 2021 CHI conference on human factors in computing sy...

  2. [10]

    Many miles to go

    Glyn Elwyn, Isabelle Scholl, Caroline Tietbohl, Mala Mann, Adrian GK Edwards, Catharine Clay, France Légaré, Trudy van der Weijden, Carmen L Lewis, Richard M Wexler, et al. 2013. “Many miles to go. . . ”: a systematic review of the implementa- tion of patient decision support ...

  3. [11]

    Robin C Feldman, Ehrik Aldana, and Kara Stein. 2019. Artificial intelligence in the health care space: how we can trust what we cannot know. Stan. L. & Pol’y Rev. 30 (2019), 399

  4. [12]

    Raymond Fok and Daniel S Weld. 2023. In search of verifiability: Explanations rarely enable complementary performance in ai-advised decision making. arXiv preprint arXiv:2305.07722 (2023)

  5. [13]

    Susanne Gaube, Harini Suresh, Martina Raue, Eva Lermer, Timo K Koch, Matthias FC Hudecek, Alun D Ackery, Samir C Grover, Joseph F Coughlin, Dieter Frey, et al. 2023. Non-task expert physicians benefit from correct explainable AI advice when reviewing X-rays. Scientific reports...

  6. [14]

    Susanne Gaube, Harini Suresh, Martina Raue, Alexander Merritt, Seth J Berkowitz, Eva Lermer, Joseph F Coughlin, John V Guttag, Errol Colak, and Marzyeh Ghas- semi. 2021. Do as AI say: susceptibility in deployment of clinical decision-aids. NPJ digital medicine 4, 1 (2021), 31

  7. [15]

    Bhavya Ghai, Q Vera Liao, Yunfeng Zhang, Rachel Bellamy, and Klaus Mueller

  8. [16]

    Marzyeh Ghassemi, Luke Oakden-Rayner, and Andrew L Beam. 2021. The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health 3, 11 (2021), e745–e750

  9. [17]

    Grace Y Gombolay, Andrew Silva, Mariah Schrum, Nakul Gopalan, Jamika Hallman-Cooper, Monideep Dutt, and Matthew Gombolay. 2024. Effects of ex- plainable artificial intelligence in neurology decision support. Annals of Clinical and Translational Neurology 11, 5 (2024), 1224–1235

  10. [18]

    Matthew Groh, Omar Badri, Roxana Daneshjou, Arash Koochek, Caleb Harris, Luis R Soenksen, P Murali Doraiswamy, and Rosalind Picard. 2024. Deep learning- aided decision support for diagnosis of skin disease across skin tones. Nature Medicine 30, 2 (2024), 573–583

  11. [19]

    Yingxuan Guo, Changke Huang, Yaying Sheng, Wenjie Zhang, Xin Ye, Hengli Lian, Jiahao Xu, and Yiqi Chen. 2024. Improve the efficiency and accuracy of ophthalmologists’ clinical decision-making based on AI technology. BMC Medical Informatics and Decision Making 24, 1 (2024), 192

  12. [20]

    Achim Hekler, Jochen S Utikal, Alexander H Enk, Axel Hauschild, Michael We- ichenthal, Roman C Maron, Carola Berking, Sebastian Haferkamp, Joachim Klode, Dirk Schadendorf, et al. 2019. Superior skin cancer classification by the combina- tion of human and artificial intelligenc...

  13. [21]

    Katharine E Henry, Rachel Kornfield, Anirudh Sridharan, Robert C Linton, Cather- ine Groh, Tony Wang, Albert Wu, Bilge Mutlu, and Suchi Saria. 2022. Human– machine teaming is key to AI adoption: clinicians’ experiences with a deployed machine learning system. NPJ digital medic...

  14. [22]

    Benjamin D Horne, Dorit Nevo, John O’Donovan, Jin-Hee Cho, and Sibel Adalı

  15. [23]

    Maia Jacobs, Melanie F Pradier, Thomas H McCoy Jr, Roy H Perlis, Finale Doshi- Velez, and Krzysztof Z Gajos. 2021. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry 11, 1 (2021), 108

  16. [24]

    Ekaterina Jussupow, Izak Benbasat, and Armin Heinzl. 2020. Why are we averse towards algorithms? A comprehensive literature review on algorithm aversion. (2020)

  17. [25]

    In Proceedings of the International AAAI Conference on Web and Social Media , Vol

    Rating reliability and bias in news articles: Does AI assistance help everyone?. In Proceedings of the International AAAI Conference on Web and Social Media , Vol. 13. 247–256

  18. [26]

    Sunnie SY Kim, Nicole Meister, Vikram V Ramaswamy, Ruth Fong, and Olga Russakovsky. 2022. HIVE: Evaluating the human interpretability of visual expla- nations. In European Conference on Computer Vision . Springer, 280–298

  19. [27]

    Why is’ Chicago’deceptive?

    Vivian Lai, Han Liu, and Chenhao Tan. 2020. " Why is’ Chicago’deceptive?" Towards Building Model-Driven Tutorials for Humans. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13

  20. [28]

    Mohsen Khosravi, Zahra Zare, Seyyed Morteza Mojtabaeian, and Reyhane Izadi

  21. [29]

    Zhongwen Li, Lei Wang, Xuefang Wu, Jiewei Jiang, Wei Qiang, He Xie, Hongjian Zhou, Shanjun Wu, Yi Shao, and Wei Chen. 2023. Artificial intelligence in oph- thalmology: The path to the real-world clinic. Cell Reports Medicine 4, 7 (2023)

  22. [30]

    Samia Massalha, Owen Clarkin, Rebecca Thornhill, Glenn Wells, and Benjamin JW Chow. 2018. Decision support tools, systems, and artificial intelligence in cardiac , Vol. 1, No. 1, Article . Publication date: December 2024. 2-Factor Retrieval for Improved Human-AI Decision Makin...

  23. [31]

    Tim Miller. 2019. Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence 267 (2019), 1–38

  24. [32]

    Vivian Lai and Chenhao Tan. 2019. On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In Proceedings of the conference on fairness, accountability, and transparency . 29–38

  25. [33]

    Mahsan Nourani, Chiradeep Roy, Jeremy E Block, Donald R Honeycutt, Tahrima Rahman, Eric Ragan, and Vibhav Gogate. 2021. Anchoring bias affects mental model formation and user reliance in explainable ai systems. In 26th International Conference on Intelligent User Interfaces . 340–350

  26. [34]

    David B Olawade, Nicholas Aderinto, Gbolahan Olatunji, Emmanuel Kokori, Aanuoluwapo C David-Olawade, and Manizha Hadi. 2024. Advancements and applications of Artificial Intelligence in cardiology: Current trends and future prospects. Journal of Medicine, Surgery, and Public He...

  27. [35]

    P Rajpurkar. 2017. CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning. ArXiv abs/1711 5225 (2017)

  28. [36]

    Mohammad Naiseh, Dena Al-Thani, Nan Jiang, and Raian Ali. 2023. How the different explanation classes impact trust calibration: The case of clinical decision support systems. International Journal of Human-Computer Studies 169 (2023), 102941

  29. [37]

    Mike Schaekermann, Graeme Beaton, Elaheh Sanoubari, Andrew Lim, Kate Larson, and Edith Law. 2020. Ambiguity-aware ai assistants for medical data analysis. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–14

  30. [38]

    Jane Scheetz, Philip Rothschild, Myra McGuinness, Xavier Hadoux, H Peter Soyer, Monika Janda, James JJ Condon, Luke Oakden-Rayner, Lyle J Palmer, Stuart Keel, et al. 2021. A survey of clinicians on the use of artificial intelligence in ophthalmology, dermatology, radiology and...

  31. [39]

    Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision . 618–626

  32. [40]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144

  33. [41]

    Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé III, and Jordan Boyd-Graber. 2023. Large Language Models Help Humans Verify Truthfulness–Except When They Are Convincingly Wrong. arXiv preprint arXiv:2310.12558 (2023)

  34. [42]

    Venkatesh Sivaraman, Leigh A Bukowski, Joel Levin, Jeremy M Kahn, and Adam Perer. 2023. Ignore, trust, or negotiate: Understanding clinician acceptance of AI-based treatment recommendations in health care. In Proceedings of the 2023 CHI Conference on Human Factors in Computing...

  35. [43]

    Nicole Sultanum, Michael Brudno, Daniel Wigdor, and Fanny Chevalier. 2018. More text please! understanding and supporting the use of visualization for clinical text overview. In Proceedings of the 2018 CHI conference on human factors in computing systems. 1–13

  36. [44]

    Soroosh Shahtalebi, S Farokh Atashzar, Rajni V Patel, Mandar S Jog, and Arash Mohammadi. 2021. A deep explainable artificial intelligent framework for neuro- logical disorders discrimination. Scientific reports 11, 1 (2021), 9630

  37. [45]

    Danielle Timmermans, Bert Molewijk, Anne Stiggelbout, and Job Kievit. 2004. Different formats for communicating surgical risks to patients and the effect on choice of treatment. Patient education and counseling 54, 3 (2004), 255–263

  38. [46]

    Philipp Tschandl, Christoph Rinner, Zoe Apalla, Giuseppe Argenziano, Noel Codella, Allan Halpern, Monika Janda, Aimilios Lallas, Caterina Longo, Josep Malvehy, et al. 2020. Human–computer collaboration for skin cancer recognition. Nature medicine 26, 8 (2020), 1229–1234

  39. [47]

    Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Ger- stenberg, Michael S Bernstein, and Ranjay Krishna. 2023. Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023), 1–38

  40. [48]

    Alan R Tait, Terri Voepel-Lewis, Brian J Zikmund-Fisher, and Angela Fagerlin

  41. [49]

    Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M Summers. 2017. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference...

  42. [50]

    Jeremy C Wyatt and Douglas G Altman. 1995. Commentary: Prognostic models: clinically useful or quickly forgotten? Bmj 311, 7019 (1995), 1539–1541

  43. [51]

    Yao Xie, Melody Chen, David Kao, Ge Gao, and Xiang’Anthony’ Chen. 2020. CheXplain: enabling physicians to explore and understand data-driven, AI-enabled medical imaging analysis. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13

  44. [52]

    Qian Yang, Aaron Steinfeld, and John Zimmerman. 2019. Unremarkable AI: Fitting intelligent decision support into critical, clinical decision-making processes. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–11

  45. [53]

    Danding Wang, Qian Yang, Ashraf Abdul, and Brian Y Lim. 2019. Designing theory- driven user-centric explainable AI. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–15

  46. [54]

    Alexandra Zytek, Dongyu Liu, Rhema Vaithianathan, and Kalyan Veeramachaneni

  47. [58]

    Feiyang Yu, Alex Moehring, Oishi Banerjee, Tobias Salz, Nikhil Agarwal, and Pranav Rajpurkar. 2024. Heterogeneity and predictors of the effects of AI assis- tance on radiologists. Nature Medicine 30, 3 (2024), 837–849

  48. [60]

    IEEE Transactions on Visualization and Computer Graphics 28, 1 (2021), 1161–1171

    Sibyl: Understanding and addressing the usability challenges of machine learning in high-stakes decision making. IEEE Transactions on Visualization and Computer Graphics 28, 1 (2021), 1161–1171. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 , Vol. 1, N...

  49. [2010]

    Journal of health communication 15, 5 (2010), 487–501

    The effect of format on parents’ understanding of the risks and benefits of clinical research: a comparison between text, tables, and graphics. Journal of health communication 15, 5 (2010), 487–501

  50. [2019]

    Hello AI

    " Hello AI": uncovering the onboarding needs of medical practitioners for human-AI collaborative decision-making. Proceedings of the ACM on Human- computer Interaction 3, CSCW (2019), 1–24

  51. [2021]

    Proceedings of the ACM on Human-Computer Interaction 4, CSCW3 (2021), 1–28

    Explainable active learning (xal) toward ai explanations as interfaces for machine teachers. Proceedings of the ACM on Human-Computer Interaction 4, CSCW3 (2021), 1–28

  52. [2024]

    Health services research and managerial epidemiology 11 (2024), 23333928241234863

    Artificial intelligence and decision-making in healthcare: a thematic anal- ysis of a systematic review of reviews. Health services research and managerial epidemiology 11 (2024), 23333928241234863

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.