REVIEW 4 major objections 5 minor 76 references
An Empirical Examination of the Evaluative AI Framework
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read In a 375-participant experiment, replacing AI recommendations with optional pro and con evidence did not improve decision quality or engagement, and users offloaded cognition much as they do with traditional AI.
desk verdict A pre-registered, powered null result for Miller's Evaluative AI framework that is honest about its own limitations; the main caveat is whether the SHAP-based evidence display actually instantiates the framework. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a five-condition randomized experiment built on the contrast between recommendation-driven and hypothesis-driven support. In the Evaluative AI condition, participants see only two buttons, one for pro evidence and one for con evidence, and can open them at any time; no recommendation is ever offered. The evidence itself is generated from SHAP feature attributions of a logistic regression model (each personal characteristic contributes a positive or negative push to the predicted probability), displayed as bar charts and converted into short textual arguments. Performance is scored with the Brier score, $\frac{1}{N}\sum_{i=1}^{N}(p_i-o_i)^2$, which also determines the bonus payment; decision time, NASA-TLX cognitive load, and qualitative coding of free-text decision descriptions complete the measurement of how users actually reasoned.
What would settle it
Re-run the same five-condition experiment with an evidence-comprehension check that participants must pass before deciding, and with the Evaluative AI group's bonus tied to using the evidence; if that group's Brier score drops significantly below the control group's, the claim that this framework does not improve decisions would be overturned in that setting.
Extended reading notes
Core claim
The central discovery is a null result, stated on the paper's own terms: in an incentivized between-subjects experiment with 375 lay participants estimating whether each of four individuals earns above-median net income, the treatment that implemented the Evaluative AI framework—pro and con evidence available on request, with no recommendation—did not improve decision-making performance. The mean Brier score $\frac{1}{N}\sum_{i=1}^{N}(p_i-o_i)^2$ for the Evaluative AI group was $0.230$, the worst among the five groups, and a Kruskal-Wallis test found no significant between-group difference ($p=0.154$). The paper also reports limited engagement: 62.64% of Evaluative AI participants clicked both evidence buttons in every round, and qualitative descriptions of the decision process showed that AI-supported participants relied on fewer features than unsupported controls, which the author interprets as cognitive offloading and potential automation bias rather than deeper reasoning with the evidence. The conclusion is that this implementation of the framework did not deliver the performance or engagement gains it was designed to produce.
Load-bearing premise
The load-bearing premise is that the pro and con evidence participants could click on was clear enough and faithful enough to what the framework intends that the experiment actually tested the framework's idea rather than a flawed presentation of it.
Editorial extensions
If this is right
- No AI support condition significantly beat doing the task alone on Brier score; the Evaluative AI condition's mean score of $0.230$ was numerically the worst, so the framework's hoped-for performance gain did not materialize.
- Evidence on demand slowed decisions: Evaluative AI (mean 51.7s), Evidence Only (56.4s), and Recommendation and Evidence (57.2s) all took significantly longer than Control or Recommendation Only (about 41s).
- Cognitive load, as measured by NASA-TLX, did not differ between any treatments, contradicting the hypothesis that hypothesis-driven support would demand more mental effort.
- Qualitative reports indicate all AI-assisted groups mentioned individual features less often than the no-AI control, which the paper reads as cognitive offloading and potential automation bias rather than deeper engagement.
- Only 62.64% of participants in the Evaluative AI condition used the optional evidence in every round, so superficial engagement persisted even when evidence was the only AI feature.
Reading between the lines
- The null result may be an implementation ceiling: if the textual pro/con evidence was hard to parse or did not match participants' mental models, the framework itself remains untested. A comprehension check on the evidence would separate these possibilities.
- The incentive design may matter: a £3 show-up fee with a bonus up to £6 per four tasks may not create the high-stakes, low-frequency setting the framework targets. A field-like task with real professional consequences could produce different engagement.
- This result fits a cost-benefit view of overreliance: when optional evidence costs extra clicks and reading time, users may rationally skip it. A design that makes evidence an unavoidable first step, or that gives users a personal stake in accuracy, could recover the intended effect.
- The binary income task may be too simple for pros-and-cons reasoning to add value; the framework's natural habitat is multi-hypothesis diagnosis where several options compete, so a multi-class extension is a direct next test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a preregistered, incentivized online experiment (N = 375) testing Miller's Evaluative AI framework. Participants estimated the probability that each of four individuals had a net income above the median, with five between-subjects conditions: no AI, AI recommendation only, evidence only, recommendation plus evidence, and the Evaluative AI condition in which pro/con evidence was available on demand. The central findings are that Brier-score performance did not differ significantly across conditions, with the Evaluative AI condition numerically worst; evidence-presenting conditions were slower than control or recommendation-only; NASA-TLX cognitive load did not differ; and qualitative coding suggested that AI-assisted participants mentioned fewer features. The authors conclude that the Evaluative AI framework, at least as implemented here, did not improve decision-making performance and that engagement with evidence was limited.
Significance. If the null result is taken at face value, it is a useful empirical counterpoint to Le et al. (2024a) and to Miller's theoretical claims, and it provides a cautionary data point for hypothesis-driven XAI. The study has real methodological strengths: it is preregistered, uses a power analysis with Bonferroni correction, applies non-parametric tests with a control group, uses monetary incentives tied to Brier score, and makes data and code available. The main weakness is that the validity of the evidence operationalization is not established, which makes the scope of the conclusion uncertain. The study is therefore worth publishing only after the authors either validate the evidence display or explicitly restrict their claims to the specific implementation tested.
major comments (4)
- [Section 3.7] Section 3.7 and Section 3.2: The pro/con evidence is generated from SHAP values and converted to free text via GPT-4o, but the manuscript provides no validation that these statements accurately reflect the SHAP contributions, no comprehensibility or manipulation check, and no evidence that participants could map the directional statements to the probability scale they had to report. The pretest mentioned in Section 1 found bar charts poorly understood, yet no check confirms that the added textual descriptions fixed the problem. Because the central null result in the Evidence Only and Evaluative AI conditions depends on participants actually receiving and understanding Miller-style evaluative evidence, the study currently cannot distinguish 'the framework does not improve decisions' from 'this operationalization of evidence was not usable.' I recommend either a comprehension check on the evidence display, an audit of the GPT-4o-generated statements, or a condition in which the evidence is coupled with the AI's baseline probability.
- [Section 3.3] Section 3.3 (Brier score) and Section 3.7: The outcome is a calibrated probability, but the evaluative evidence is presented as feature-level SHAP contributions without any stated mapping from those log-odds-style contributions to the percentage estimates participants must produce. Evidence Only and Evaluative AI participants never see the AI's probability estimate or a base-rate anchor, whereas Recommendation Only participants do. This confounds the presence of a recommendation with the availability of a response-scale cue; the observed advantage of Recommendation Only and the poor numerical performance of Evaluative AI may reflect the absence of a probability anchor rather than the absence of a recommendation. The paper should either provide the model's predicted probability or an explicit aggregation instruction in the evidence conditions, or restrict the conclusion to the specific evidence presentation tested.
- [Section 3.4] Section 3.4 (Decision-Making Process) and Section 4.4: The qualitative coding is described as performed by the author with the assistance of GPT-4o, but no inter-rater reliability, double-coding subsample, or validation of the LLM-assisted coding is reported. The claims about AI mentions, feature mentions, and the resulting cognitive-offloading interpretation rest on these categories. A second human coder on a random subsample with a kappa statistic, or a documented validation of the GPT-4o coding protocol, is necessary to support those conclusions.
- [Abstract and Section 4.5] The abstract characterizes engagement as 'limited,' but Section 4.5 reports that 62.64% of Evaluative AI participants clicked on both the pro and con evidence in all four rounds (i.e., all eight evidence items). If the intended claim is that participants did not deeply process the evidence, the click data alone do not support that; the paper should state the specific measure on which 'limited engagement' is based and reconcile the high click-through rate with the superficial-engagement interpretation.
minor comments (5)
- [Section 3.5] There is a typo in the subsection heading ('Desicion-Making Process'); similarly, 'Apendix' appears in Section 3.2.
- [Section 4.1] Reporting effect sizes and confidence intervals for the pairwise comparisons (or at least the test statistics) would improve interpretability, since Kruskal-Wallis p-values alone do not convey the magnitude of the null result.
- [Appendix A] Appendix A states the AI is correct 77% of the time, whereas Section 3.7 reports 15/20 (75%) accuracy with a 50% threshold; these numbers should be reconciled.
- [Section 3.8] The final group sizes are noticeably uneven (62 to 91 participants); the manuscript should state how randomization was implemented and whether the imbalance was checked.
- [Section 4.4] The pairwise chi-squared tests are reported only through p-values; the test statistics or exact p-values with the correction method should be given.
Circularity Check
No significant circularity: the paper reports an empirical null result; its claims are observed outcomes, not derivations forced by definitions or self-citation.
full rationale
This is an empirical, pre-registered behavioral experiment, not a derivation chain. The central claim—that the Evaluative AI implementation did not improve decision performance and saw limited evidence engagement—is a statistical finding from independent dependent variables (Brier score, decision time, NASA-TLX, click behavior). The hypotheses H1–H3 follow from Miller's framework, but the outcome measures are not defined in terms of the framework's assumptions, so the hypotheses could have been confirmed or refuted by the data; indeed they were refuted. The pro/con evidence is an operationalization of the framework via SHAP signs and GPT-4o text, and the paper's Limitations explicitly concede that this presentation may need improvement and that the model did not greatly outperform participants. Those are construct-validity and generalizability concerns, not circularity: the study does not define 'Evaluative AI improved performance' to mean 'the SHAP/GPT-4o display was engaged with,' nor does it fit any parameter and then relabel that fit as a prediction. There is no self-citation chain that carries the argument: the only self-citation (Kornowicz and Thommes, 2024) appears in Limitations as a supporting reference for domain dependence and is not load-bearing. No uniqueness theorem, ansatz, or renaming of a known result is used to make the conclusion true by construction. The null result is an observed difference among experimental conditions, not an identity or a reduction to the study's inputs. Therefore no circular step is identifiable, and the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The SHAP-based pro/con evidence, converted to text by GPT-4o, is a faithful and comprehensible operationalization of the evidence that Miller's Evaluative AI framework envisions.
- domain assumption The income-estimation task with 20 features and lay participants is an appropriate testbed for the framework, which targets medium/high-stakes decisions.
- domain assumption Participants' self-reports of their decision processes (coded by the author with GPT-4o) can be treated as valid evidence about cognitive engagement.
Cite this review
Pith. "Pith review of An Empirical Examination of the Evaluative AI Framework." pith.science (2026). https://pith.science/paper/UBWAZHO7
@misc{pith2026241108583,
author = {Pith},
title = {Pith review of: An Empirical Examination of the Evaluative AI Framework},
year = {2026},
howpublished = {\url{https://pith.science/paper/UBWAZHO7}},
note = {Machine review of arXiv:2411.08583}
}
read the original abstract
This study empirically examines the "Evaluative AI" framework, which aims to enhance the decision-making process for AI users by transitioning from a recommendation-based approach to a hypothesis-driven one. Rather than offering direct recommendations, this framework presents users pro and con evidence for hypotheses to support more informed decisions. However, findings from the current behavioral experiment reveal no significant improvement in decision-making performance and limited user engagement with the evidence provided, resulting in cognitive processes similar to those observed in traditional AI systems. Despite these results, the framework still holds promise for further exploration in future research.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...
-
[2]
Albrecht, J. P. (2016). How the gdpr will change the world. Eur. Data Prot. L. Rev. , 2:287
work page 2016
-
[3]
Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., and Weld, D. S. (2021). Does the whole exceed its parts? the effect of ai explanations on complementary team performance. (arXiv:2006.14779). arXiv:2006.14779 [cs]
arXiv 2021
-
[4]
Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garcia, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R., and Herrera, F. (2020). Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information Fusion , 58:82–115
work page 2020
-
[5]
Becker, B. and Kohavi, R. (1996). Adult . UCI Machine Learning Repository. DOI : https://doi.org/10.24432/C5XW20
doi:10.24432/c5xw20 1996
-
[6]
Bertrand, A., Viard, T., Belloum, R., Eagan, J. R., and Maxwell, W. (2023). On selective, mutable and dialogic xai: a review of what users say about different types of interactive explanations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , page 1–21, Hamburg Germany. ACM
work page 2023
-
[7]
Bogard, J. and Shu, S. (2022). Algorithm Aversion and the Aversion to Counter-Normative Decision Procedures
work page 2022
-
[8]
Buçinca, Z., Malaya, M. B., and Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction , 5(CSCW1):188:1--188:21
work page 2021
Show all 76 references
-
[9]
E., Doshi-Velez, F., and Gajos, K
Buçinca, Z., Swaroop, S., Paluch, A. E., Doshi-Velez, F., and Gajos, K. Z. (2024). Contrastive explanations that anticipate human misconceptions can improve human decision-making skills. (arXiv:2410.04253). arXiv:2410.04253
2024 arXiv
-
[10]
Carton, S., Mei, Q., and Resnick, P. (2020). Feature-based explanations don’t help people detect misclassifications of online toxicity. Proceedings of the International AAAI Conference on Web and Social Media , 14:95–106
2020
-
[11]
Castelnovo, A., Crupi, R., Mombelli, N., Nanino, G., and Regoli, D. (2023). Evaluative item-contrastive explanations in rankings. (arXiv:2312.10094). arXiv:2312.10094 [cs]
2023 arXiv
-
[12]
W., and Lehmann, D
Castelo, N., Bos, M. W., and Lehmann, D. R. (2019). Task-dependent algorithm aversion. Journal of Marketing Research , 56(5):809–825
2019
-
[13]
L., Schonger, M., and Wickens, C
Chen, D. L., Schonger, M., and Wickens, C. (2016). otree—an open-source platform for laboratory, online, and field experiments. Journal of Behavioral and Experimental Finance , 9:88--97
2016
-
[14]
V., Vaughan, J
Chen, V., Liao, Q. V., Vaughan, J. W., and Bansal, G. (2023). Understanding the role of human intuition on reliance in human-ai decision-making with explanations. (arXiv:2301.07255). arXiv:2301.07255 [cs]
2023 arXiv
-
[15]
M., and Zhu, H
Cheng, H.-F., Wang, R., Zhang, Z., O’Connell, F., Gray, T., Harper, F. M., and Zhu, H. (2019). Explaining decision-making algorithms through ui: Strategies to help non-expert stakeholders. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , CHI ’1...
2019
-
[16]
Chew, R., Bollenbacher, J., Wenger, M., Speer, J., and Kim, A. (2023). Llm-assisted content analysis: Using large language models to support deductive coding. (arXiv:2306.14924). arXiv:2306.14924
2023 arXiv
-
[17]
Chromik, M., Eiband, M., Buchner, F., Krüger, A., and Butz, A. (2021). I think i get your point, ai! the illusion of explanatory depth in explainable ai. In 26th International Conference on Intelligent User Interfaces , IUI ’21, page 307–317, New York, NY, USA. Association for...
2021
-
[18]
C., Sui, Y., Kumar, B., and Vouitsis, N
Cresswell, J. C., Sui, Y., Kumar, B., and Vouitsis, N. (2024). Conformal prediction sets improve human decision making. (arXiv:2401.13744). arXiv:2401.13744 [cs, stat]
2024 arXiv
-
[19]
Croskerry, P. (2009). A universal model of diagnostic reasoning. Academic medicine , 84(8):1022--1028
2009
-
[20]
Faul, F., Erdfelder, E., Lang, A.-G., and Buchner, A. (2007). G* power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior research methods , 39(2):175--191
2007
-
[21]
Gajos, K. Z. and Mamykina, L. (2022). Do people engage cognitively with ai? impact of ai assistance on incidental learning. In 27th International Conference on Intelligent User Interfaces , page 794–806, Helsinki Finland. ACM
2022
-
[22]
V., Zhang, Y., Bellamy, R., and Mueller, K
Ghai, B., Liao, Q. V., Zhang, Y., Bellamy, R., and Mueller, K. (2020). Explainable active learning (xal): An empirical study of how local explanations impact annotator experience. (arXiv:2001.09219). arXiv:2001.09219 [cs]
2020 arXiv
-
[23]
o der, C., and Schupp, J. (2019). The german socio-economic panel (soep). Jahrb \
Goebel, J., Grabka, M. M., Liebig, S., Kroh, M., Richter, D., Schr \"o der, C., and Schupp, J. (2019). The german socio-economic panel (soep). Jahrb \"u cher f \"u r National \"o konomie und Statistik , 239(2):345--360
2019
-
[24]
S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A
Goh, E., Gallo, R., Hom, J., Strong, E., Weng, Y., Kerman, H., Cool, J., Kanjee, Z., Parsons, A. S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A. P., Rodman, A., and Chen, J. H. (2024). Influence of a large language model on diagnostic reasoning: A randomized clin...
2024
-
[25]
Gouveia, S. S. and Malík, J. (2024). Crossing the trust gap in medical ai: Building an abductive bridge for xai. Philosophy & Technology , 37(3):105
2024
-
[26]
Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., and Pedreschi, D. (2018). A survey of methods for explaining black box models. ACM Computing Surveys , 51(5):93:1--93:42
2018
-
[27]
Hart, S. G. (2006). Nasa-task load index (nasa-tlx); 20 years later. Proceedings of the Human Factors and Ergonomics Society Annual Meeting , 50(9):904–908
2006
-
[28]
Hemmer, P., Schemmer, M., Kühl, N., Vössing, M., and Satzger, G. (2024). Complementarity in human-ai collaboration: Concept, sources, and evidence. (arXiv:2404.00029). arXiv:2404.00029 [cs]
2024 arXiv
-
[29]
Hemmer, P., Schemmer, M., Vössing, M., and Kühl, N. (2021). Human-ai complementarity in hybrid intelligence systems: A structured literature review. In PACIS 2021 Proceedings
2021
-
[30]
Herm, L.-V. (2023). Impact of explainable ai on cognitive load: Insights from an empirical study
2023
-
[31]
Jacovi, A., Marasović, A., Miller, T., and Goldberg, Y. (2021). Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in ai. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , FAccT ’21, page 624–635...
2021
-
[32]
Kaur, H., Nori, H., Jenkins, S., Caruana, R., Wallach, H., and Wortman Vaughan, J. (2020). Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning. In Proceedings of the 2020 CHI Conference on Human Factors in Computing ...
2020
-
[33]
K., Rall, E
Klein, G., Phillips, J. K., Rall, E. L., and Peluso, D. A. (2007). A data--frame theory of sensemaking. In Expertise out of context , pages 118--160. Psychology Press
2007
-
[34]
and Thommes, K
Kornowicz, J. and Thommes, K. (2024). Algorithm, expert, or both? evaluating the role of feature selection methods on user preferences and reliance. (arXiv:2408.01171). arXiv:2408.01171 [cs]
2024 arXiv
-
[35]
V., and Tan, C
Lai, V., Chen, C., Smith-Renner, A., Liao, Q. V., and Tan, C. (2023a). Towards a science of human-ai decision making: An overview of design space in empirical human-subject studies. In 2023 ACM Conference on Fairness, Accountability, and Transparency , page 1369–1385, Chicago ...
2023
-
[36]
why is ‘chicago’ deceptive?
Lai, V., Liu, H., and Tan, C. (2020). “why is ‘chicago’ deceptive?” towards building model-driven tutorials for humans. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , page 1–13, Honolulu HI USA. ACM
2020
-
[37]
and Tan, C
Lai, V. and Tan, C. (2019). On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In Proceedings of the Conference on Fairness, Accountability, and Transparency , page 29–38, Atlanta GA USA. ACM
2019
-
[38]
V., and Tan, C
Lai, V., Zhang, Y., Chen, C., Liao, Q. V., and Tan, C. (2023b). Selective explanations: Leveraging human input to align explainable ai. (arXiv:2301.09656). arXiv:2301.09656 [cs]
2023 arXiv
-
[39]
Le, T., Miller, T., Singh, R., and Sonenberg, L. (2023). Explaining model confidence using counterfactuals. Proceedings of the AAAI Conference on Artificial Intelligence , 37(1010):11856–11864
2023
-
[40]
Le, T., Miller, T., Sonenberg, L., and Singh, R. (2024a). Towards the new xai: A hypothesis-driven approach to decision support using evidence. (arXiv:2402.01292). arXiv:2402.01292 [cs]
2024 arXiv
-
[41]
Le, T., Miller, T., Zhang, R., Sonenberg, L., and Singh, R. (2024b). Visual evaluative ai: A hypothesis-driven tool with concept-based explanations and weight of evidence. (arXiv:2407.04710). arXiv:2407.04710 [cs]
2024 arXiv
-
[42]
Liu, H., Lai, V., and Tan, C. (2021). Understanding the effect of out-of-distribution examples and interactive explanations on human-ai decision making. Proceedings of the ACM on Human-Computer Interaction , 5(CSCW2):408:1--408:45
2021
-
[43]
Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R., editors, Advances in Neural Information Processing Systems 30 , pages 4765--4774. ...
2017
-
[44]
and Coiera, E
Lyell, D. and Coiera, E. (2017). Automation bias and verification complexity: a systematic review. Journal of the American Medical Informatics Association , 24(2):423–431
2017
-
[45]
Ma, S., Lei, Y., Wang, X., Zheng, C., Shi, C., Yin, M., and Ma, X. (2023). Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision-making. In Proceedings of the 2023 CHI Conference on Human Factors i...
2023
-
[46]
MacCarthy, M. (2019). An examination of the algorithmic accountability act of 2019. Available at SSRN 3615731
2019
-
[47]
Mahmud, H., Islam, A. K. M. N., Ahmed, S. I., and Smolander, K. (2022). What influences algorithmic decision-making? a systematic literature review on algorithm aversion. Technological Forecasting and Social Change , 175:121390
2022
-
[48]
Malone, T., Vaccaro, M., Campero, A., Song, J., Wen, H., and Almaatouq, A. (2023). A test for evaluating performance in human-ai systems
2023
-
[49]
Mayring, P. (2015). Qualitative Content Analysis: Theoretical Background and Procedures , page 365–380. Springer Netherlands, Dordrecht
2015
-
[50]
Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence , 267:1–38
2019
-
[51]
Miller, T. (2023). Explainable ai is dead, long live explainable ai!: Hypothesis-driven decision support using evaluative ai. In 2023 ACM Conference on Fairness, Accountability, and Transparency , page 333–342, Chicago IL USA. ACM
2023
-
[52]
A., Veness, J., Bellemare, M
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015). Human-le...
2015
-
[53]
M., Carignan, D., and Horvitz, E
Nori, H., King, N., McKinney, S. M., Carignan, D., and Horvitz, E. (2023). Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375
2023 arXiv
-
[54]
Peirce, C. S. (2009). Writings of Charles S. Peirce: a chronological edition, volume 8: 1890--1892 , volume 8. Indiana University Press
2009
-
[55]
Popper, K. (2014). Conjectures and refutations: The growth of scientific knowledge . routledge
2014
-
[56]
G., Hofman, J
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J. W., and Wallach, H. (2021). Manipulating and measuring model interpretability. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , CHI ’21, page 1–52, New York, NY, USA. A...
2021
-
[57]
T., Singh, S., and Guestrin, C
Ribeiro, M. T., Singh, S., and Guestrin, C. (2018). Anchors: High-precision model-agnostic explanations. Proceedings of the AAAI Conference on Artificial Intelligence , 32(11)
2018
-
[58]
Risko, E. F. and Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences , 20(9):676–688
2016
-
[59]
Rogha, M. (2023). Explain to decide: A human-centric review on the role of explainable artificial intelligence in ai-assisted decision making. (arXiv:2312.11507). arXiv:2312.11507 [cs]
2023 arXiv
-
[60]
Rong, Y., Leemann, T., Nguyen, T.-t., Fiedler, L., Seidel, T., Kasneci, G., and Kasneci, E. (2022). Towards human-centered explainable ai: User studies for model explanations. (arXiv:2210.11584). arXiv:2210.11584 [cs]
2022 arXiv
-
[61]
Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence , 1(5):206–215
2019
-
[62]
Sawyer, S. F. (2009). Analysis of variance: The fundamental concepts. Journal of Manual & Manipulative Therapy
2009
-
[63]
Schemmer, M., Hemmer, P., Nitsche, M., Kühl, N., and Vössing, M. (2022). A meta-analysis of the utility of explainable artificial intelligence in human-ai decision-making. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society , page 617–626. arXiv:2205.05126 [cs]
2022 arXiv
-
[64]
Schemmer, M., Kuehl, N., Benz, C., Bartos, A., and Satzger, G. (2023). Appropriate reliance on ai advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces , page 410–422, Sydney NSW Australia. ACM
2023
-
[65]
Schuff, D., Corral, K., and Turetken, O. (2011). Comparing the understandability of alternative data warehouse schemas: An empirical study. Decision Support Systems , 52(1):9–20
2011
-
[66]
A., Scheidegger, C., and Roy, C
Slack, D., Friedler, S. A., Scheidegger, C., and Roy, C. D. (2019). Assessing the local interpretability of machine learning models. (arXiv:1902.03501). arXiv:1902.03501
2019 arXiv
-
[67]
Spatola, N. (2024). The efficiency-accountability tradeoff in ai integration: Effects on human performance and over-reliance. Computers in Human Behavior: Artificial Humans , page 100099
2024
-
[68]
H., Bentley, L
Tai, R. H., Bentley, L. R., Xia, X., Sitt, J. M., Fankhauser, S. C., Chicas-Mosier, A. M., and Monteith, B. G. (2024). An examination of the use of large language models to aid analysis of textual data. International Journal of Qualitative Methods , 23:16094069241231168
2024
-
[69]
S., and Krishna, R
Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., and Krishna, R. (2023). Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction , 7(CSCW1):1–38
2023
-
[70]
Vered, M., Livni, T., Howe, P. D. L., Miller, T., and Sonenberg, L. (2023). The effects of explanations on automation bias. Artificial Intelligence , 322:103952
2023
-
[71]
Wang, D., Yang, Q., Abdul, A., and Lim, B. Y. (2019). Designing theory-driven user-centric explainable ai. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , page 1–15, Glasgow Scotland Uk. ACM
2019
-
[72]
and Yin, M
Wang, X. and Yin, M. (2021). Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making. In 26th International Conference on Intelligent User Interfaces , IUI ’21, page 318–328, New York, NY, USA. Association for Computing Machinery
2021
-
[73]
Wischnewski, M., Krämer, N., and Müller, E. (2023). Measuring and understanding trust calibrations for automated systems: A survey of the state-of-the-art and future directions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , CHI ’23, page 1–1...
2023
-
[74]
Yates, J. F. and Potworowski, G. A. (2012). Evidence-based decision management
2012
-
[75]
L., and Li, X
You, S., Yang, C. L., and Li, X. (2022). Algorithmic versus human advice: Does presenting prediction performance matter for algorithm appreciation? Journal of Management Information Systems , 39(2):336–365
2022
-
[76]
V., and Bellamy, R
Zhang, Y., Liao, Q. V., and Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency , FAT* ’20, page 295–305, New York, ...
2020
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.