Pith. sign in

REVIEW 3 major objections 5 minor 61 references

Can Domain Experts Rely on AI Appropriately? A Case Study on AI-Assisted Prostate Cancer MRI Diagnosis

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper reports that radiologists reading prostate MRI with AI help beat their own unaided accuracy but still trail the AI alone because they override it too often, and that majority-voting across AI-assisted radiologists can surpass…

desk verdict Solid expert-level study confirming under-reliance in prostate MRI, but the flashy ensemble claim is not supported by the analysis as presented. read the letter →

arxiv 2502.03482 v1 pith:EHIL3BLM submitted 2025-02-03 eess.IV cs.AIcs.CVcs.CYcs.HCcs.LG

classification eess.IVcs.AIcs.CVcs.CYcs.HCcs.LG
keywords human-AIdecisionmakingunder-reliancedomainexpertsprostatecancerMRIradiologistperformanceAIassistancemajority-voteensemblefeedback
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether trained radiologists use AI assistance appropriately when diagnosing prostate cancer from MRI scans, and whether knowing their own performance makes them use it better. It reports that AI-assisted radiologists outperform unaided radiologists, but their team performance still falls short of the AI alone, a gap the authors attribute to under-reliance: doctors override the AI's advice on a sizable share of disagreements, often incorrectly. Giving radiologists performance feedback and showing the AI prediction up front made them follow the AI more often, but did not significantly improve overall accuracy. The most forward-looking result is that a simple majority vote across eight AI-assisted radiologists can outperform the AI alone, suggesting that collective human-AI decision making can be complementary even when individual assisted readers are not.

What carries the argument

The experimental engine is a three-condition comparison with practicing radiologists: an independent human read, an AI-assisted read made after seeing the AI's prediction and lesion annotations, and an AI-first read made after performance feedback. The behavioral mechanism measured is reliance: how often a radiologist's final decision agrees with the AI, how often they override it, and whether the override is correct. The ensemble analysis uses a majority vote over the eight radiologists' final predictions, with ties broken by reported confidence. Under-reliance is defined by the gap between the AI-alone decision and the human-AI final decision: when the radiologist and AI disagree, the radiologist keeps an incorrect independent answer often enough to drag team performance below AI-alone.

What would settle it

Run the same two workflows in a 2x2 design: with and without performance feedback, and with the AI shown before or after the radiologist's own diagnosis. If a feedback-only arm changes reliance or accuracy as much as the up-front-AI arm does, then the paper's conclusion that performance feedback did not significantly improve human-AI teams would be wrong.

Watch

Extended reading notes

Core claim

The central claim is that expert human-AI teams in prostate MRI diagnosis occupy a middle band: they are more accurate than the same radiologists working alone, but less accurate than the model by itself because of under-reliance. In Study 1, radiologists made an independent diagnosis, then saw the AI prediction, then finalized their decision; in Study 2, after a memory washout and after receiving feedback on their own, the AI's, and their team's previous performance, they saw the AI prediction before diagnosing. Neither workflow produced team accuracy that reached the AI's standalone performance on the paired common-case comparisons, and aggregate performance feedback did not significantly improve the team. Yet when the eight radiologists' AI-assisted final decisions were combined by majority vote, the ensemble outperformed the AI on AUROC, accuracy, and precision. The paper interprets this as evidence that complementary performance is achievable, but not at the level of the individual assisted reader.

Load-bearing premise

The feedback conclusion rests on comparing Study 1 with Study 2, but Study 2 changed both the feedback and the timing of AI (shown up front, with no independent diagnosis) at the same time, with no feedback-only control arm, so the effect of feedback alone is not isolated.

Editorial extensions

If this is right

  • Deploying AI as a second reader in prostate MRI improves expert accuracy, but the full accuracy of the AI is not realized unless radiologists follow its advice more closely.
  • Showing AI predictions before the human commits to a diagnosis increases adoption of AI advice and can improve sensitivity, but by itself does not close the gap to AI-alone performance.
  • Aggregating multiple AI-assisted reads by majority vote is a promising deployment pattern, since the ensemble can outperform both the average radiologist and the AI alone.
  • Performance feedback about one's own and the AI's accuracy does not, in this setting, significantly improve human-AI team accuracy.
  • The behavioral pattern of under-reliance seen in crowdworker studies also appears among board-certified domain experts, suggesting the finding is not an artifact of lay participants.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's claims, if under-reliance is structural rather than informational, interventions such as requiring radiologists to justify overrides or defaulting to the AI recommendation may be more effective than performance feedback.
  • Beyond the paper's claims, the ensemble gains could grow if the radiologists were deliberately diverse in experience or in prostate-zonal expertise, because the error patterns of individual readers would be less correlated.
  • Beyond the paper's claims, the finding that humans are better at catching the AI's false positives than its missed cases suggests a division-of-labor routing strategy: let the AI flag negatives for human confirmation, while escalating AI-positive cases for extra review.
  • Beyond the paper's claims, a clinical deployment would need a cost-benefit comparison between AI alone and a human-AI ensemble, since the ensemble requires multiple expert reads per case and the AI model itself is inexpensive to run.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports two pre-registered human studies with eight practicing radiologists on AI-assisted prostate cancer MRI diagnosis. Study 1 asks radiologists to give an independent diagnosis, then show the AI prediction, then finalize; Study 2 provides individual performance feedback from Study 1 and shows the AI prediction up front without an independent diagnosis. The authors report that human-AI teams outperform humans alone, still underperform the AI alone, that performance feedback did not significantly improve team performance, that upfront AI increases AI adoption, and that a majority vote of human-AI teams can outperform the AI alone.

Significance. The study is valuable because it brings real domain experts into the human-AI decision-making loop, uses biopsy-confirmed labels, and pre-registers the protocol. The behavioral measurements (agreement, follow/overrule rates, confidence, time) are directly observed and bootstrap testing is applied throughout. If the stated effects hold, they strengthen the empirical basis for claims about under-reliance in expert populations and point to ensemble aggregation as a potentially practical deployment strategy. The main weaknesses are the confounded Study 2 design and the lack of out-of-sample validation for the ensemble claim, both of which are load-bearing for central conclusions, so the current manuscript needs revision before the claims are fully supported.

major comments (3)
  1. [§3.4, §4.2] The conclusion that performance feedback did not lead to significant improvements is not supported by the design, because Study 2 changes both the feedback and the AI presentation timing simultaneously. Participants in Study 2 receive performance feedback and also see the AI prediction up front without making an independent diagnosis, whereas Study 1 requires an initial diagnosis and shows the AI afterward. Any observed null effect on accuracy or any change in AI adoption could be due to either manipulation or their interaction, so the paper should not attribute the result to feedback specifically. A control arm with upfront AI but no feedback, or a factorial design, is needed to disentangle these factors.
  2. [§4.1, Table 3, Appendix E] The claim that the majority vote of human-AI teams can outperform AI alone is based on a single test set with shared cases, and the bootstrapped p-values are computed on the same cases used to form the ensemble. Appendix E (Table 13) shows that on the common-50 subset the effect weakens substantially (e.g., Study 1 H+AI ensemble vs AI AUROC p=0.265, accuracy p=0.216; Study 2 AUROC p=0.112, accuracy p=0.229), which indicates the result is highly dependent on the particular case mix. Without a holdout set, repeated split-half evaluation, or an independent set of radiologists/cases, the paper cannot support the general statement that ensemble voting outperforms AI alone.
  3. [§4.1, Table 1, Table 2] The statement that human-AI teams 'consistently outperform humans alone' overstates the evidence in Study 2, where the human-alone baseline is taken from Study 1 on a different case set and the common-case paired analysis shows non-significant differences in several metrics (e.g., AUROC p=0.074, specificity p=0.450 in Study 2 common-50). The direction is consistent, but the wording should be qualified to reflect which metrics are significant and which are not, especially since multiple metrics are tested without adjustment.
minor comments (5)
  1. [§3.3] The interface description mentions 'BWI' as one of the image sequences, while the rest of the paper and the dataset description refer to 'DWI' (diffusion-weighted imaging); this inconsistency should be fixed.
  2. [Global] The running header on the first page reads 'Trovato and Tobin, et al.' which appears to be a template artifact and should be replaced with the manuscript's own title and authors.
  3. [§4.2] The labels in Fig. 6a are somewhat hard to parse; the text describing the four sub-groups would benefit from a clearer mapping between the diagram boxes and the accuracy values in the text.
  4. [§3.5] The statistical testing section does not state whether any correction for multiple comparisons was applied across the many metrics and conditions; even if no correction is intended, it should be stated explicitly so readers can calibrate the reported p-values.
  5. [Appendix B] In the demographics description, 'whilte' should be 'while'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: all central claims are direct empirical measurements against independent biopsy-confirmed test labels.

full rationale

This paper contains no derivation chain that reduces to its own inputs. Every headline quantity — human-alone accuracy, human+AI accuracy, AI-alone accuracy, reliance rates, and ensemble performance — is an empirical measurement computed from recorded radiologist decisions, AI model outputs, and biopsy-confirmed PI-CAI test labels, with bootstrap z-tests on those measurements. The 'under-reliance' conclusion is supported by an independent behavioral statistic (radiologists change their initial decision in only 20.4% of disagreement cases, while the follow-AI subgroup has higher accuracy), not by definitional equivalence. The ensemble claim is a majority-vote aggregation of the same measured decisions and is therefore an empirical comparison, not a fitted parameter renamed as a prediction. The only self-citations (Lai et al. and Lai and Tan) appear as related-work context and are not load-bearing for any result. Potential concerns about the feedback condition being confounded with AI timing, and about the ensemble result resting on a single split, are validity and robustness issues rather than circularity, because no target quantity is assumed in order to produce it.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No new entities are postulated, and the central claim is an empirical measurement rather than a derivation. The conclusions rely on domain assumptions about biopsy ground truth, washout, confidence-based AUROC, and bootstrap inference, plus hand-set thresholds for AI prediction and lesion overlap. No fitted free parameters in the classical sense appear in the human-AI analysis itself.

free parameters (3)
  • AI binary prediction threshold = 0.5 (lesion-level)
    Hand-set threshold converts nnU-Net lesion segmentation into binary case-level AI predictions; accuracy, sensitivity, and specificity comparisons depend on it, while AUROC does not.
  • Lesion-level overlap threshold = 10%
    Lesion hits and misses are defined by 10% overlap between predicted and ground-truth annotations; changing this threshold changes the per-lesion accuracy numbers.
  • Ensemble tie-break weighting
    In 4-4 majority-vote ties, the ensemble uses radiologists' reported confidence; this hand-chosen rule affects ensemble outcomes in tied cases.
assumptions (5)
  • domain assumption Testing-case labels from biopsy are accurate ground truth.
    Section 3.1 states all testing cases are biopsy-confirmed; the paper uses these as the reference for all metrics.
  • domain assumption The 30-day minimum washout removes memory and learning effects between studies.
    Section 3.4 states the washout period is used to eliminate any recall effects, but no empirical check of washout is provided.
  • domain assumption Radiologists' confidence-slider values can be used as a continuous score for AUROC.
    Section 3.5 reports AUROC for humans, which requires a ranking score; the paper does not explicitly describe how human AUROC is computed from the five-level confidence slider.
  • standard math Bootstrap z-tests with resampling provide valid inference for N=8 and 75/100 cases.
    Section 3.5 uses 10,000 bootstrap iterations; the validity of the normal approximation with 8 participants and clustered cases is assumed.
  • domain assumption The AI's binary predictions and lesion maps are a fixed reference intervention.
    The paper does not vary AI threshold or calibration across participants; comparisons treat the AI as a static assistive tool.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can Domain Experts Rely on AI Appropriately? A Case Study on AI-Assisted Prostate Cancer MRI Diagnosis." pith.science (2026). https://pith.science/paper/EHIL3BLM

@misc{pith2026250203482,
  author       = {Pith},
  title        = {Pith review of: Can Domain Experts Rely on AI Appropriately? A Case Study on AI-Assisted Prostate Cancer MRI Diagnosis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EHIL3BLM}},
  note         = {Machine review of arXiv:2502.03482}
}
read the original abstract

Despite the growing interest in human-AI decision making, experimental studies with domain experts remain rare, largely due to the complexity of working with domain experts and the challenges in setting up realistic experiments. In this work, we conduct an in-depth collaboration with radiologists in prostate cancer diagnosis based on MRI images. Building on existing tools for teaching prostate cancer diagnosis, we develop an interface and conduct two experiments to study how AI assistance and performance feedback shape the decision making of domain experts. In Study 1, clinicians were asked to provide an initial diagnosis (human), then view the AI's prediction, and subsequently finalize their decision (human-AI team). In Study 2 (after a memory wash-out period), the same participants first received aggregated performance statistics from Study 1, specifically their own performance, the AI's performance, and their human-AI team performance, and then directly viewed the AI's prediction before making their diagnosis (i.e., no independent initial diagnosis). These two workflows represent realistic ways that clinical AI tools might be used in practice, where the second study simulates a scenario where doctors can adjust their reliance and trust on AI based on prior performance feedback. Our findings show that, while human-AI teams consistently outperform humans alone, they still underperform the AI due to under-reliance, similar to prior studies with crowdworkers. Providing clinicians with performance feedback did not significantly improve the performance of human-AI teams, although showing AI decisions in advance nudges people to follow AI more. Meanwhile, we observe that the ensemble of human-AI teams can outperform AI alone, suggesting promising directions for human-AI collaboration.

Figures

Figures reproduced from arXiv: 2502.03482 by the authors.

Figure 1
Figure 1. Overview of our experiments with radiologists. In study 1, participant radiologists (N=8) reviewed 75 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Screenshots of the webapp interface for our human study. (a) Fig. 2a presents a user interface for [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. An example of lesion-level annotation comparing human experts (red contour), AI (yellow contour), [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Individual radiologists performance compared with the AI model. The model achieves higher per [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Mean performance of Human-alone, Human+AI, Human-ensemble, Human+AI-ensemble, and AI in [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Comparison of Human-AI Decision Alignment and Accuracy. Blue shading indicates frequency of [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: The average confidence score, time spent, sensitivity, and NPV on the common 50-case subset for [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Login page [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: Consent page [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: Toy demonstration example page [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: Exit survey for study 1 [PITH_FULL_IMAGE:figures/full_fig_p023_11.png]
Figure 12
Figure 12. Figure 12: Exit survey for study 2 [PITH_FULL_IMAGE:figures/full_fig_p024_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 57 canonical work pages

  1. [1]

    Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar, and Tobias Salz. 2023. Combining human expertise with artificial intelligence: Experimental evidence from radiology . Technical Report. National Bureau of Economic Research

  2. [2]

    automation bias

    Saar Alon-Barkat and Madalina Busuioc. 2023. Human–AI interactions in public sector decision making:“automation bias” and “selective adherence” to algorithmic advice. Journal of Public Administration Research and Theory 33, 1 (2023), 153–169

  3. [3]

    Lucrezia Greta Armando, Gianluca Miglio, Pierluigi de Cosmo, and Clara Cena. 2023. Clinical decision support systems to improve drug prescription and therapy optimisation in clinical practice: a scoping review. BMJ Health & Care Informatics 30, 1 (2023)

  4. [4]

    Reuben Binns, Max Van Kleek, Michael Veale, Ulrik Lyngs, Jun Zhao, and Nigel Shadbolt. 2018. ’It’s Reducing a Human Being to a Percentage’ Perceptions of Justice in Algorithmic Decisions. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–14

  5. [5]

    Joeran S Bosma, Anindo Saha, Matin Hosseinzadeh, Ilse Slootweg, Maarten de Rooij, and Henkjan Huisman. 2021. Annotation-efficient cancer detection with report-guided lesion annotation for deep learning-based prostate cancer detection in bpMRI. arXiv preprint arXiv:2112.05151 (2021)

  6. [6]

    Aritrick Chatterjee, Ambereen N Yousuf, Roger Engelmann, Carla Harmath, Grace Lee, Milica Medved, Ernest B Jamison, Abel Lorente Campos, Batuhan Gundogdu, Glenn Gerber, et al. 2025. Prospective Validation of an Automated Hybrid Multidimensional MRI Tool for Prostate Cancer Detection Using Targeted Biopsy: Comparison with PI-RADS- based Assessment. Radiolo...

  7. [7]

    Maarten De Rooij, Esther HJ Hamoen, Jurgen J Fütterer, Jelle O Barentsz, and Maroeska M Rovers. 2014. Accuracy of multiparametric MRI for prostate cancer detection: a meta-analysis. American Journal of Roentgenology 202, 2 (2014), 343–351

  8. [8]

    Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2015. Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General 144, 1 (2015), 114

Show all 61 references
  1. [9]

    Ben Green and Yiling Chen. 2019. Disparate interactions: An algorithm-in-the-loop analysis of fairness in risk assessments. In Proceedings of the conference on fairness, accountability, and transparency . 90–99

  2. [10]

    H Benjamin Harvey and Vrushab Gowda. 2020. How the FDA regulates AI. Academic radiology 27, 1 (2020), 58–61

  3. [11]

    Ahmed Hosny, Chintan Parmar, John Quackenbush, Lawrence H Schwartz, and Hugo JWL Aerts. 2018. Artificial intelligence in radiology. Nature Reviews Cancer 18, 8 (2018), 500–510

  4. [12]

    Fabian Isensee, Paul F Jaeger, Simon AA Kohl, Jens Petersen, and Klaus H Maier-Hein. 2021. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18, 2 (2021), 203–211

  5. [13]

    Ayush Jain, David Way, Vishakha Gupta, Yi Gao, Guilherme de Oliveira Marinho, Jay Hartford, Rory Sayres, Kimberly Kanada, Clara Eng, Kunal Nagpal, et al. 2021. Development and assessment of an artificial intelligence–based tool for skin condition diagnosis by primary care phys...

  6. [14]

    Amirhossein Kiani, Bora Uyumazturk, Pranav Rajpurkar, Alex Wang, Rebecca Gao, Erik Jones, Yifan Yu, Curtis P Langlotz, Robyn L Ball, Thomas J Montine, et al. 2020. Impact of a deep learning assistant on the histopathologic classification of liver cancer. NPJ digital medicine 3...

  7. [15]

    Hyo-Eun Kim, Hak Hee Kim, Boo-Kyung Han, Ki Hwan Kim, Kyunghwa Han, Hyeonseob Nam, Eun Hye Lee, and Eun- Kyung Kim. 2020. Changes in cancer detection and false-positive recall in mammography using artificial intelligence: a retrospective, multireader study. The Lancet Digital ...

  8. [16]

    Jon Kleinberg, Himabindu Lakkaraju, Jure Leskovec, Jens Ludwig, and Sendhil Mullainathan. 2018. Human decisions and machine predictions. The quarterly journal of economics 133, 1 (2018), 237–293

  9. [17]

    Marie-Luise Kromrey, Laura Steiner, Felix Schön, Julie Gamain, Christian Roller, and Carolin Malsch. 2024. Navigating the Spectrum: Assessing the Concordance of ML-Based AI Findings with Radiology in Chest X-Rays in Clinical Settings. In Healthcare, Vol. 12. MDPI, 2225

  10. [18]

    Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. 2019. An evaluation of the human-interpretability of explanation. arXiv preprint arXiv:1902.00006 (2019)

  11. [19]

    Vivian Lai, Chacha Chen, Q Vera Liao, Alison Smith-Renner, and Chenhao Tan. 2023. Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies. In Proceedings of FAccT

  12. [20]

    Vivian Lai and Chenhao Tan. 2019. On human predictions with explanations and predictions of machine learning mod- els: A case study on deception detection. In Proceedings of the Conference on Fairness, Accountability, and Transparency . 29–38

  13. [21]

    Himabindu Lakkaraju, Stephen H Bach, and Jure Leskovec. 2016. Interpretable decision sets: A joint framework for description and prediction. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1675–1684

  14. [22]

    Curtis P Langlotz. 2019. Will artificial intelligence replace radiologists? , e190058 pages

  15. [23]

    Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, et al. 2023. Towards accurate differential diagnosis with large language models. arXiv preprint arXiv:2312.00164 (2023)

  16. [24]

    Scott Mayer McKinney, Marcin Sieniek, Varun Godbole, Jonathan Godwin, Natasha Antropova, Hutan Ashrafian, Trevor Back, Mary Chesus, Greg S Corrado, Ara Darzi, et al. 2020. International evaluation of an AI system for breast cancer screening. Nature 577, 7788 (2020), 89–94

  17. [25]

    Justin G Norden and Nirav R Shah. 2022. What AI in health care can learn from the long road to autonomous vehicles. NEJM Catalyst Innovations in Care Delivery 3, 2 (2022)

  18. [26]

    Khaled Ouanes and Nesren Farhah. 2024. Effectiveness of artificial intelligence (AI) in clinical decision support systems and care delivery. Journal of Medical Systems 48, 1 (2024), 74

  19. [27]

    Allison Park, Chris Chute, Pranav Rajpurkar, Joe Lou, Robyn L Ball, Katie Shpanskaya, Rashad Jabarkheel, Lily H Kim, Emily McKenna, Joe Tseng, et al. 2019. Deep learning–assisted diagnosis of cerebral aneurysms using the HeadXNet model. JAMA network open 2, 6 (2019), e195600–e195600

  20. [28]

    Ruben Pauwels, Danieli Moura Brasil, Mayra Cristina Yamasaki, Reinhilde Jacobs, Hilde Bosmans, Deborah Queiroz Freitas, and Francisco Haiter-Neto. 2021. Artificial intelligence for detection of periapical lesions on intraoral radiographs: Comparison between convolutional neura...

  21. [29]

    Pranav Rajpurkar, Chloe O’Connell, Amit Schechter, Nishit Asnani, Jason Li, Amirhossein Kiani, Robyn L Ball, Marc Mendelson, Gary Maartens, Daniël J van Hoving, et al. 2020. CheXaid: deep learning assistance for physician diagnosis of tuberculosis using chest x-rays in patient...

  22. [30]

    Andreas M Rauschecker, Jeffrey D Rudie, Long Xie, Jiancong Wang, Michael Tran Duong, Emmanuel J Botzolakis, Asha M Kovalovich, John Egan, Tessa C Cook, R Nick Bryan, et al. 2020. Artificial intelligence system approaching neuroradiologist-level differential diagnosis accuracy ...

  23. [31]

    Carlo Reverberi, Tommaso Rigon, Aldo Solari, Cesare Hassan, Paolo Cherubini, and Andrea Cherubini. 2022. Ex- perimental evidence of effective human–AI collaboration in medical decision-making. Scientific reports 12, 1 (2022), 14952

  24. [32]

    Alejandro Rodriguez-Ruiz, Kristina Lång, Albert Gubern-Merida, Mireille Broeders, Gisella Gennaro, Paola Clauser, Thomas H Helbich, Margarita Chevalier, Tao Tan, Thomas Mertelmeier, et al. 2019. Stand-alone artificial intelligence for breast cancer detection in mammography: co...

  25. [33]

    Noordman, Ivan Slootweg, Christian Roest, Stefan J

    Anindo Saha, Joeran S Bosma, Jasper J Twilt, Bram van Ginneken, Anders Bjartell, Anwar R Padhani, David Bonekamp, Geert Villeirs, Georg Salomon, Gianluca Giannarini, Jayashree Kalpathy-Cramer, Jelle Barentsz, Klaus H Maier-Hein, Mirabela Rusu, Olivier Rouvière, Roderick van de...

  26. [34]

    Maxi Scherer. 2019. Artificial Intelligence and Legal Decision-Making: The Wide Open? Journal of international arbitration 36, 5 (2019)

  27. [35]

    Jarrel CY Seah, Cyril HM Tang, Quinlan D Buchlak, Xavier G Holt, Jeffrey B Wardman, Anuar Aimoldin, Nazanin Esmaili, Hassan Ahmad, Hung Pham, John F Lambert, et al. 2021. Effect of a comprehensive deep-learning model on the accuracy of chest x-ray interpretation by radiologist...

  28. [36]

    Yongsik Sim, Myung Jin Chung, Elmar Kotter, Sehyo Yune, Myeongchan Kim, Synho Do, Kyunghwa Han, Hanmyoung Kim, Seungwook Yang, Dong-Jae Lee, et al . 2020. Deep convolutional neural network–based software improves radiologist detection of malignant lung nodules on chest radiogr...

  29. [37]

    David F Steiner, Robert MacDonald, Yun Liu, Peter Truszkowski, Jason D Hipp, Christopher Gammage, Florence Thng, Lily Peng, and Martin C Stumpe. 2018. Impact of deep learning assistance on the histopathologic review of lymph nodes for metastatic breast cancer. The American jou...

  30. [38]

    Nan Wu, Jason Phang, Jungkyu Park, Yiqiu Shen, Zhe Huang, Masha Zorin, Stanisław Jastrzębski, Thibault Févry, Joe Katsnelson, Eric Kim, et al. 2019. Deep neural networks improve radiologists’ performance in breast cancer screening. IEEE transactions on medical imaging 39, 4 (2...

  31. [39]

    Annotate Cancer

    Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020. Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-Assisted Decision Making. arXiv preprint arXiv:2001.02114 (2020). A Model Impementation Details Training configurationsWe use the established n...

  32. [40]

    Please select your current role in the medical field: Resident Fellow Attending Physician Other (Please specify):

  33. [41]

    How would you rate your level of experience with prostate MRI? Novice (I have little to no experience) Intermediate (I have moderate experience and have interpreted a few cases) Advanced (I am very experienced and regularly perform/interpret prostate MRI) Expert (I possess spe...

  34. [42]

    Where do you practice? Academic Medical Center Community Hospital Private Practice Other (Please specify):

  35. [43]

    Country of Practice:

  36. [44]

    Gender: Male Female Non-binary/third gender Prefer not to say Prefer to self-describe: Section 2: Opinions on AI

  37. [46]

    How accurate do you believe the AI's predictions were?

  38. [47]

    How useful was the AI in identifying lesion areas for you?

  39. [49]

    During the task involving AI, to what extent did you feel stressed, insecure, discouraged, irritated, or annoyed?

  40. [51]

    What improvements would you suggest for the AI tool? Section 3: Final Comments Please share any additional comments or insights you have about using AI in medical diagnostics. Not familiar at all Somewhat unfamiliar Neutral Somewhat familiar Very familiar Not accurate at all S...

  41. [52]

    How helpful did you find the performance feedback from the first stage of the study?

  42. [53]

    Strongly agree Agree Neutral Disagree Strongly disagree

    Rate the following statement: The performance feedback on AI and human accuracy, sensitivity, and specificity affects your trust in the AI system. Strongly agree Agree Neutral Disagree Strongly disagree

  43. [54]

    How did the information about team performance influence your approach to working with the AI? Encouraged more collaboration No change in approach Discouraged collaboration Other (please specify):

  44. [55]

    Strongly agree Agree Neutral Disagree Strongly disagree

    Rate the following statement: Your prior experience with AI improved your performance in this phase. Strongly agree Agree Neutral Disagree Strongly disagree

  45. [56]

    How would you rate the overall collaboration experience with the AI in this phase compared to the first phase? Much better Better About the same Worse Much worse Section 2: Opinions on AI

  46. [57]

    How familiar are you with AI technology in medicine?

  47. [58]

    How accurate do you believe the AI's predictions were in this study?

  48. [59]

    How useful was the AI in identifying lesion areas for you in this study?

  49. [60]

    Would you trust an AI's predictions in your daily practice?

  50. [61]

    During the task involving AI, to what extent did you feel stressed, insecure, discouraged, irritated, or annoyed in this phase?

  51. [62]

    After this experience, how likely are you to consider using AI assistance in your future clinical practice?

  52. [63]

    Did the AI-assisted predictions influence your diagnostic decisions? If yes, how?

  53. [64]

    What improvements would you suggest for the AI tool? Section 3: Final Comments Please share any additional comments or insights you have about using AI in medical diagnostics. Not helpful at all Slightly helpful Moderately helpful Very helpful Extremely helpful Not familiar at...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.