Pith. sign in

REVIEW 2 major objections 6 minor 2 cited by

Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

T0 review · 2 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Explanations increase reliance on LLM answers—correct or not—while sources and contradictions curb overreliance on wrong ones.

desk verdict Strong pre-registered study with a credible sources effect; the inconsistency sub-claim is real but currently confounded by question identity. read the letter →

arxiv 2502.08554 v1 pith:YDA3ZISB submitted 2025-02-12 cs.HC cs.AI

classification cs.HCcs.AI
keywords largelanguagemodelsoverrelianceappropriaterelianceexplanationssourcesinconsistencieshuman-AIinteractionquestionanswering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Everyday users of large language models face a hard question: when should a fluent, confident answer be trusted? The paper tackles this through two empirical studies—a think-aloud session with 16 users and a pre-registered experiment with 308 participants—and claims that three response features govern reliance: explanations (supporting details), inconsistencies inside explanations, and clickable sources. Its central finding is that explanations increase reliance on both correct and incorrect answers, while sources improve reliance selectively (more agreement with correct answers, less with incorrect ones) and inconsistent explanations reduce overreliance on incorrect answers. If these findings hold, interface designers have concrete levers—provide accurate sources, surface contradictions—to help users get the benefit of LLM assistance without blind deference.

What carries the argument

The load-bearing apparatus is a 2 × 2 × 2 within-subjects experiment using 12 difficult binary factual questions, where each of 308 participants saw eight response types from a hypothetical LLM named Theta: {correct, incorrect} × {no explanation, explanation} × {no sources, clickable sources}. Explanations—supporting details that justify the answer—and answers were generated in advance with ChatGPT and Perplexity AI so that content could be controlled, and reliance was measured behaviorally as whether the participant's final answer agreed with Theta's answer, complemented by self-reported confidence, justification-quality and actionability ratings, source-clicking, and follow-up questions. In addition, the authors coded naturally occurring inconsistencies in the explanations (sets of statements that cannot both be true) and ran a pre-registered ANOVA comparing incorrect answers with no explanation, consistent explanation, and inconsistent explanation; a 16-person think-aloud study supplied the qualitative account of how users notice these cues.

What would settle it

Run the same 12 questions with matched explanation pairs that are identical except for one internal contradiction and randomly assign participants to versions; if agreement with incorrect answers does not drop when the contradiction is present, the paper's inconsistency claim is falsified.

Watch

Extended reading notes

Core claim

Users agree with an LLM's answer more often when the answer comes with an explanation, regardless of whether the answer is actually right; the paper shows this in a controlled setting where the same difficult questions are paired with correct or incorrect answers, with or without explanations and sources. Sources change the pattern: they raise agreement when the answer is correct and lower it when the answer is wrong, and they increase time on task and the odds of overriding an incorrect answer. Explanations that contain a logical inconsistency—for instance, an answer that contradicts the numbers cited to support it—produce significantly less agreement and higher accuracy than consistent explanations when the answer is wrong. The paper therefore concludes that explanation is not a single good thing: its effect depends on whether it invites verification (sources) and whether it contains visible cracks (inconsistencies).

Load-bearing premise

The inconsistency finding rests on comparing a small set of questions whose explanations happened to contain contradictions against the other questions' explanations, so the result holds only if the contradiction itself—not the content or difficulty of those three questions—is what changed participants' agreement and accuracy.

Editorial extensions

If this is right

  • Explanations alone make users more likely to accept an answer, so adding them without other safeguards can actively increase overreliance on wrong LLM outputs.
  • Providing clickable, accurate sources is a concrete corrective: in the no-explanation condition it raised agreement with correct answers from 67.2% to 73.4% and lowered agreement with incorrect answers from 78.2% to 68.2%.
  • When the LLM answer is wrong, sources without explanation produce the highest user accuracy (31.8%), while explanation alone gives the lowest (17.2%); when the answer is right, explanation plus sources gives the highest accuracy (79.9%).
  • Inconsistent explanations cut overreliance: agreement with wrong answers fell from 83.3% to 69.7% and accuracy rose from 16.7% to 30.3% compared with consistent explanations.
  • Because the beneficial source effect was obtained with real, mostly accurate links, the authors expect that fake, broken, or irrelevant sources would not help and could even increase perceived credibility.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the causal story is that inconsistency triggers deeper scrutiny, then automatically detecting and highlighting contradictions (for example, by checking whether the answer matches the numbers in the explanation) should reproduce the effect without requiring users to spot the flaw themselves.
  • Beyond the paper: the paper's observational comparison suggests a testable design rule—LLM answers whose supporting explanation is internally consistent should be treated as more reliable for answer selection, since consistency of the explanation is itself predictive of correctness in the paper's stimulus set.
  • Beyond the paper: the source effect may depend on individual differences: users who click links (roughly 119 of 308 participants clicked in at least one task, while 189 never clicked) may be the main beneficiaries, so an interface that actively nudges source-checking could widen the benefit to the majority who do not click.
  • Beyond the paper: the experiment used deliberately hard questions where lay users have little prior knowledge, so the findings may not transfer to domains where users can independently evaluate the answer; testing with familiar topics would show whether sources still carry the same corrective force.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper investigates how features of LLM responses shape users' reliance when answering objective questions. Study 1 is a think-aloud study (N=16) that identifies explanations, inconsistencies, and sources as key features. Study 2 is a pre-registered, within-subjects experiment (N=308) with a 2x2x2 design (answer correctness, presence of explanation, presence of clickable sources) using realistic LLM-generated responses from ChatGPT and Perplexity AI. The main mixed-effects analyses show that explanations increase agreement with both correct and incorrect answers, while sources increase appropriate reliance on correct answers and reduce overreliance on incorrect answers. A secondary observational analysis suggests that inconsistent explanations are associated with reduced overreliance on incorrect answers. The authors discuss implications for designing LLM interfaces to foster appropriate reliance.

Significance. The randomized manipulation of explanations and sources, combined with mixed-effects models that include participant and question random effects, makes the main 2x2x2 findings credible and directly relevant to the HCI community. The pre-registration, power analysis, and use of realistic LLM-generated stimuli are notable strengths that increase confidence in the explanation and source effects. If the inconsistency finding were causally supported, the design recommendation to highlight inconsistencies would be of considerable practical value; however, as it stands, that sub-claim is not yet established because the analysis is observational and confounded.

major comments (2)
  1. [§4.3.1, Figure 5] The comparison of consistent vs. inconsistent explanations is confounded by question identity. Only 3 of the 12 task questions produced naturally occurring inconsistent explanations, all for incorrect answers, and these three questions may differ systematically in difficulty, answer plausibility, or specific content (e.g., the arithmetic error in the Brazil population item). The ANOVA does not include participant or question random effects, and it collapses across the sources manipulation, so the reported differences (agreement 69.7% vs. 83.3%; accuracy 30.3% vs. 16.7%) cannot be attributed to inconsistency per se. This confound is load-bearing because the abstract presents the inconsistency result as a finding and §5.1 uses it to recommend interventions that highlight inconsistencies.
  2. [Abstract and §5.1] The causal claim that inconsistencies reduce overreliance, and the design recommendation to highlight inconsistencies as an intervention, are stronger than the evidence supports. The independent variable was not manipulated; it was observed post hoc after the experiment. Even a mixed-effects reanalysis would address non-independence but would not address selection on question content or the small number of inconsistent items. The authors should either reframe the inconsistency result as an exploratory, hypothesis-generating finding or conduct a follow-up experiment that factorially manipulates inconsistency while holding content constant.
minor comments (6)
  1. [§4.2.2] In the sentence reporting the confidence effect, β = .96, SE = .10, p < .001, the formatting of the p-value is inconsistent with the rest of the paper; please standardize the formatting of regression results throughout.
  2. [§4.3.1] The text does not state the cell size for the inconsistent explanation condition (N=155); adding this number would help readers interpret the precision of the estimates.
  3. [Figure 5] The y-axis labels are not fully specified in the figure caption; please indicate the response scale or unit for each panel to improve readability.
  4. [§5.3] The limitations section does not mention the confound in the inconsistency analysis; given the prominence of this finding, it should be explicitly acknowledged as an observational analysis with limited internal validity.
  5. [§4.1.4] The coding of inconsistencies was performed by the authors without reporting inter-rater reliability; since this variable is central to the inconsistency sub-claim, a second coder or a reliability statistic would strengthen confidence in the coding.
  6. [Appendix A] There is a typo in the first sentence: the the relationship should be the relationship.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's central claims are tested in a pre-registered randomized experiment, and the observational inconsistency analysis raises internal-validity concerns but does not reduce to its own inputs.

full rationale

The paper's derivation chain is empirical rather than definitional. Study 1 (think-aloud) is used only to generate hypotheses about explanations, inconsistencies, and sources; Study 2 then tests those hypotheses in a pre-registered 2x2x2 within-subjects experiment in which explanation presence and source presence are randomly manipulated. The main reliance and accuracy findings are estimated from mixed-effects models with participant and question random effects, so they are not forced by construction. The inconsistency analysis in Section 4.3.1 is observational: the authors explicitly state that 'the presence of inconsistencies is not something we control for or manipulate,' and they compare naturally occurring inconsistent explanations against consistent explanations. This creates a potential confound with question identity and content, which is a validity limitation that the paper partially acknowledges, but it is not circular reasoning because the inconsistency variable is defined independently of the outcome (as 'sets of statements that cannot be true at the same time') and the observed effect could plausibly have gone in either direction. The paper's self-citations (e.g., prior work on uncertainty expression and interpretability) are used for methodological precedent and literature positioning, not as load-bearing proof of the current findings. No fitted parameter is renamed as a prediction, and no result is equivalent to its inputs by definition.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rely on standard experimental assumptions plus two domain-specific premises: the single-response setup generalizes to real interactions, and the source effect is conditional on high-quality sources. The inconsistency sub-claim additionally depends on the reliability of the author's coding and, critically, on the assumption that the 3 questions with naturally inconsistent explanations isolate the effect of inconsistency rather than question content.

assumptions (4)
  • domain assumption Participants' behavior in a single-response controlled experiment reflects how users rely on LLMs in multi-turn interactions.
    Section 4.1.1 shows exactly one controlled response per task, and Section 5.3 acknowledges this limitation.
  • domain assumption The sources used in Study 2 are real, relevant, and accurate, so the effect of sources may depend on source quality.
    Section 4.1.4 states all sources were real and relevant; Section 5.1 discusses that fake or irrelevant sources might not help.
  • standard math Mixed-effects regression assumptions, such as linearity and normality of residuals, hold for the fitted models.
    Section 4.1.3 specifies logistic and linear mixed-effects models without reporting diagnostic checks.
  • domain assumption The coding of inconsistencies by the first author is accurate and reproducible.
    Section 4.1.4 describes the coding process but no inter-rater reliability is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies." pith.science (2026). https://pith.science/paper/YDA3ZISB

@misc{pith2026250208554,
  author       = {Pith},
  title        = {Pith review of: Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YDA3ZISB}},
  note         = {Machine review of arXiv:2502.08554}
}
read the original abstract

Large language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mitigating such overreliance is a key challenge. Through a think-aloud study in which participants use an LLM-infused application to answer objective questions, we identify several features of LLM responses that shape users' reliance: explanations (supporting details for answers), inconsistencies in explanations, and sources. Through a large-scale, pre-registered, controlled experiment (N=308), we isolate and study the effects of these features on users' reliance, accuracy, and other measures. We find that the presence of explanations increases reliance on both correct and incorrect responses. However, we observe less reliance on incorrect responses when sources are provided or when explanations exhibit inconsistencies. We discuss the implications of these findings for fostering appropriate reliance on LLMs.

Figures

Figures reproduced from arXiv: 2502.08554 by the authors.

Figure 1
Figure 1. Overview of our studies. In Study 1, participants engaged in multi-turn interactions with ChatGPT to arrive at correct [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Screenshots of Study 2’s experimental task. Here the LLM response provides an incorrect answer, includes sources, [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Types of LLM responses used in Study 2. We vary three variables in the LLM responses: accuracy of the LLM’s answer to [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Summary of participants’ accuracy in Study 2. We plot the raw data means and 95% confidence intervals for participants’ [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Study 2 results on inconsistencies. We plot the raw data means and 95% confidence intervals. Brackets indicate [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Humans overrely on overconfident language models, across languages

    cs.CL 2025-07 conditional novelty 7.0 of 10

    LLMs produce overconfident-sounding answers in all five tested languages, and bilingual users show the highest overreliance risk in Japanese despite its frequent hedges.

  2. ContextBuddy: AI-Enhanced Contextual Insights for Security Alert Investigation (Applied to Intrusion Detection)

    cs.CR 2025-06 conditional novelty 6.0 of 10

    ContextBuddy trains an imitation-learning assistant on RL-simulated analysts' context requests and shows its suggestions improve alert classification accuracy and speed in simulation and a small non-expert user study.

Reference graph

Works this paper leans on

147 extracted references · 19 canonical work pages · cited by 2 Pith papers

  1. [1]

    Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. 2024. Faithful- ness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models. arXiv:cs.CL/2402.04614 https://arxiv.org/abs/2402.04614

  2. [2]

    McFarlane

    Hussam Alkaissi and Samy I. McFarlane. 2023. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 15, 2 (2023). https://doi. org/10.7759/cureus.35179

  3. [3]

    Ravinithesh Annapureddy, Alessandro Fornaroli, and Daniel Gatica-Perez. 2024. Generative AI Literacy: Twelve Defining Competencies. Digit. Gov.: Res. Pract. (aug 2024). https://doi.org/10.1145/3685680 Just Accepted

  4. [4]

    Sara Aronowitz and Tania Lombrozo. 2020. Experiential Explanation. Topics in Cognitive Science 12, 4 (2020), 1321–1336. https://doi.org/10.1111/tops.12445 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/tops.12445

  5. [5]

    Pepa Atanasova, Oana-Maria Camburu, Christina Lioma, Thomas Lukasiewicz, Jakob Grue Simonsen, and Isabelle Augenstein. 2023. Faithfulness Tests for Natural Language Explanations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . Association for Computational Linguistics, Toronto, Canada, ...

  6. [6]

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New Yor...

  7. [7]

    Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Ben- netot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2020. Ex- plainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Informatio...

  8. [8]

    Christos Bechlivanidis, David A Lagnado, Jeffrey C Zemla, and Steven Sloman

Show all 147 references
  1. [9]

    Boren and J

    T. Boren and J. Ramey. 2000. Thinking Aloud: Reconciling Theory and Practice. IEEE Transactions on Professional Communication 43, 3 (2000), 261–278. https: //doi.org/10.1109/47.867942

  2. [10]

    Richard E Boyatzis. 1998. Transforming Qualitative Information: Thematic Anal- ysis and Code Development . sage

  3. [11]

    Gelman Brandy N

    Susan A. Gelman Brandy N. Frazier and Henry M. Wellman. 2016. Young Children Prefer and Remember Satisfying Explanations. Journal of Cognition and Development 17, 5 (2016), 718–736. https://doi.org/10.1080/15248372.2015. 1098649 PMID: 28713222

  4. [12]

    Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psy- chology. Qualitative Research in Psychology 3, 2 (2006), 77–101. https: //doi.org/10.1191/1478088706qp063oa

  5. [13]

    Sylvain Bromberger. 1966. Why-Questions. In Readings in the Philosophy of Science, Baruch A. Brody (Ed.). Prentice Hall, Inc., Englewood Cliffs, 66–84

  6. [14]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 188 (apr 2021), 21 pages. https://doi.org/10.1145/3449287

  7. [15]

    Ben Buchanan, Andrew Lohn, Micah Musser, and Katerina Sedova. 2021. Truth, Lies, and Automation: How Language Models Could Change Disinformation . Re- port. Center for Security and Emerging Technology. https://doi.org/10.51593/ 2021CA003

  8. [16]

    Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020. Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems. In Proceedings of the 25th International Conference on Intelligent User Interfaces. 454–464

  9. [17]

    Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems. In Proceedings of the 2015 International Conference on Healthcare Informatics (ICHI ’15). IEEE Computer Society, USA, 160–169. https...

  10. [18]

    Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry. 2021. Onboarding Materials as Cross-functional Boundary Objects for Developing AI Assistants. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems (CHI EA ’21) . A...

  11. [19]

    Shiye Cao and Chien-Ming Huang. 2022. Understanding User Reliance on AI in Assisted Decision-Making. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 471 (nov 2022), 23 pages. https://doi.org/10.1145/3555572

  12. [20]

    Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal

    Valerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal. 2023. Understanding the Role of Human Intuition on Reliance in Human-AI Decision- Making with Explanations. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 370 (oct 2023). https://doi.org/10.1145/3610219

  13. [21]

    Cheng-Han Chiang and Hung-yi Lee. 2024. Over-Reasoning and Redundant Calculation of Large Language Models. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) , Yvette Graham and Matthew Purver...

  14. [22]

    Leah Chong, Guanglu Zhang, Kosa Goucher-Lambert, Kenneth Kotovsky, and Jonathan Cagan. 2022. Human Confidence in Artificial Intelligence and in Themselves: The Evolution and Impact of Confidence on Adoption of AI Advice. Computers in Human Behavior 127 (2022), 107018. https://...

  15. [23]

    Christiano, Jan Leike, Tom B

    Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep Reinforcement Learning from Human Preferences. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17). Curran Associates Inc., R...

  16. [24]

    Olanubi, Joseph M

    Michelle Cohn, Mahima Pushkarna, Gbolahan O. Olanubi, Joseph M. Moran, Daniel Padgett, Zion Mengesha, and Courtney Heldreth. 2024. Believing Anthro- pomorphism: Examining the Role of Anthropomorphic Cues on Trust in Large Language Models. In Extended Abstracts of the CHI Confe...

  17. [25]

    Collins, Albert Q

    Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik. 2024. Evaluating Language Models for Mathematics through In...

  18. [26]

    Francisco Cruz and Tania Lombrozo. 2024. The Effect of Jargon on Perceptions of Explanation Quality: Reconciling Contradictory Findings. In Proceedings of the Annual Meeting of the Cognitive Science Society , Vol. 46. https://escholarship. org/uc/item/4ds9s5tj

  19. [27]

    Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models.Journal of Legal Analysis 16, 1 (06 2024), 64–93. https://doi.org/10.1093/jla/laae003

  20. [28]

    Rafferty, and Christopher D

    Marie-Catherine de Marneffe, Anna N. Rafferty, and Christopher D. Manning

  21. [29]

    Igor Douven and Patricia Mirabile. 2018. Best, Second-Best, and Good-Enough Explanations: How They Matter to Reasoning. Journal of Experimental Psy- chology: Learning, Memory, and Cognition 44, 11 (2018), 1792–1813. https: //doi.org/10.1037/xlm0000545

  22. [31]

    Malin Eiband, Daniel Buschek, Alexander Kremer, and Heinrich Hussmann

  23. [32]

    Hovy, Hinrich Schütze, and Yoav Goldberg

    Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and Improving Consistency in Pretrained Language Models. Transactions of the Association for Computational Linguistics 9 (2021), 1012–1031. h...

  24. [33]

    Raymond Fok and Daniel S. Weld. [n.d.]. In Search of Verifiability: Explanations Rarely Enable Complementary Performance in AI-advised Decision Making. AI Magazine ([n. d.]). https://doi.org/10.1002/aaai.12182

  25. [34]

    Mark C Fox, K Anders Ericsson, and Ryan Best. 2011. Do Procedures for Verbal Reporting of Thinking Have to be Reactive? A Meta-Analysis and Recommenda- tions for Best Reporting Methods. Psychological Bulletin 137, 2 (2011), 316–344. https://doi.org/10.1037/a0021663

  26. [35]

    Bas. C. van Fraassen. 1980. The Scientific Image . Oxford University Press. https://doi.org/10.1093/0198244274.001.0001

  27. [36]

    Frazier, Susan A

    Brandy N. Frazier, Susan A. Gelman, and Henry M. Wellman. 2009. Preschoolers’ Search for Explanatory Information Within Adult–Child Conversation. Child Development 80, 6 (2009), 1592–1611. https://doi.org/10.1111/j.1467-8624.2009. 01356.x

  28. [37]

    Gajos and Lena Mamykina

    Krzysztof Z. Gajos and Lena Mamykina. 2022. Do People Engage Cognitively with AI? Impact of AI Assistance on Incidental Learning. In Proceedings of the 27th International Conference on Intelligent User Interfaces (IUI ’22) . Association for Computing Machinery, New York, NY, U...

  29. [38]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:cs.CL/2312.10997 https://arxiv.org/abs/2312.10997

  30. [39]

    Carly Giffin, Daniel Wilkenfeld, and Tania Lombrozo. 2017. The Explanatory Effect of a Label: Explanations with Named Categories are More Satisfying. Cognition 168 (2017), 357–369. https://doi.org/10.1016/j.cognition.2017.07.011

  31. [40]

    Ana Valeria González, Gagan Bansal, Angela Fan, Yashar Mehdad, Robin Jia, and Srinivasan Iyer. 2021. Do Explanations Help Users Detect Errors in Open- Domain QA? An Evaluation of Spoken vs. Visual Explanations. InFindings of the Association for Computational Linguistics: ACL-I...

  32. [41]

    Ben Green and Yiling Chen. 2019. The Principles and Limits of Algorithm-in- the-Loop Decision Making. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 50 (nov 2019), 24 pages. https://doi.org/10.1145/3359152

  33. [42]

    Peter Green and Catriona J. MacLeod. 2016. SIMR: An R package for Power Analysis of Generalized Linear Mixed Models by Simulation. Methods in Ecology and Evolution 7, 4 (2016), 493–498. https://doi.org/10.1111/2041-210X.12504

  34. [43]

    Gaole He, Stefan Buijsman, and Ujwal Gadiraju. 2023. How Stated Accuracy of an AI System and Analogies to Explain Accuracy Affect Human Reliance on the System. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 276 (oct 2023), 29 pages. https://doi.org/10.1145/3610067

  35. [44]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset. In Proceedings of the Neural Information Pro- cessing Systems Track on Datasets and Benchmar...

  36. [45]

    Hopkins, Deena Skolnick Weisberg, and Jordan C.V

    Emily J. Hopkins, Deena Skolnick Weisberg, and Jordan C.V. Taylor. 2019. Does Expertise Moderate the Seductive Allure of Reductive Explanations? Acta Psychologica 198 (2019), 102890. https://doi.org/10.1016/j.actpsy.2019.102890

  37. [46]

    Hopkins, Deena S

    Emily J. Hopkins, Deena S. Weisberg, and Jordan C. V. Taylor. 2016. The Se- ductive Allure is a Reductive Allure: People Prefer Scientific Explanations that Contain Logically Irrelevant Reductive Information. Cognition 155 (2016), 67–76. https://doi.org/10.1016/j.cognition.2016.06.011

  38. [47]

    Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards Reasoning in Large Language Models: A Survey. In Findings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toro...

  39. [48]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. arXiv:cs.CL/2311....

  40. [49]

    Patrick J. Hurley. 2000. A Concise Introduction to Logic . Wadsworth, Belmont, CA

  41. [51]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 55, 12, Article 248 (mar 2023), 38 pages. https://doi.org/10.1145/3571730

  42. [52]

    Daniel Kahneman. 2003. A Perspective on Judgment and Choice: Mapping Bounded Rationality. The American Psychologist 58, 9 (2003), 697–720. https: //doi.org/10.1037/0003-066X.58.9.697

  43. [53]

    Daniel Kahneman. 2011. Thinking, Fast and Slow . Farrar, Straus and Giroux

  44. [54]

    Ho, Percy Liang, and Arvind Narayanan

    Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston, Stella Biderman, Miranda Bogen, Rumman Chowdhury, Alex Engler, Peter Henderson, Yacine Jer- nite, Seth Lazar, Stefano Maffulli, Alondra Nelson, Joelle Pi...

  45. [55]

    Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020. Interpreting Interpretability: Under- standing Data Scientists’ Use of Interpretability Tools for Machine Learning. In Proceedings of the 2020 CHI Conference on Huma...

  46. [56]

    Frank C. Keil. 2006. Explanation and Understanding. Annual Review of Psychol- ogy 57 (2006), 227–254. https://doi.org/10.1146/annurev.psych.57.102904.190100

  47. [57]

    Deborah Kelemen, Joshua Rottman, and Rebecca Seston. 2013. Professional Physical Scientists Display Tenacious Teleological Tendencies: Purpose-based Reasoning as a Cognitive Default. Journal of Experimental Psychology: General 142, 4 (2013), 1074–1083. https://doi.org/10.1037/a0030399

  48. [58]

    National Geographic Kids. 2017. Weird But True! Human Body: 300 Outrageous Facts about Your A wesome Anatomy. National Geographic. https://books.google. com/books?id=dJo8DgAAQBAJ

  49. [59]

    2023.Weird But True World 2024

    National Geographic Kids. 2023.Weird But True World 2024. National Geographic. https://books.google.com/books?id=JMiPzwEACAAJ

  50. [60]

    I’m Not Sure, But

    Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. "I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust. In Proceedings of the 2024 ACM Conference on Fa...

  51. [61]

    Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong, and Olga Russakovsky. 2022. HIVE: Evaluating the Human Interpretability of Visual Explanations. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part...

  52. [62]

    Help Me Help the AI

    Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (...

  53. [63]

    Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. Humans, AI, and Context: Understanding End-Users’ Trust in a Real-World Computer Vision Application. In Proceedings of the 2023 ACM Conference on Fairness, Accountability,...

  54. [64]

    Yoonsu Kim, Jueon Lee, Seoyoung Kim, Jaehyuk Park, and Juho Kim. 2024. Un- derstanding Users’ Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level. InProceedings of the 29th International Conference on Intelligent User Interfaces ...

  55. [65]

    Kurkul and Kathleen H

    Katelyn E. Kurkul and Kathleen H. Corriveau. 2018. Question, Explanation, Follow-Up: A Mechanism for Learning From Others? Child Development 89, 1 (2018), 280–294. https://doi.org/10.1111/cdev.12726

  56. [66]

    Philippe Laban, Lidiya Murakhovs’ka, Caiming Xiong, and Chien-Sheng Wu

  57. [67]

    Bennett, and Marti A

    Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Incon- sistency Detection in Summarization. Transactions of the Associ- ation for Computational Linguistics 10 (02 2022), 163–177. https: //doi.org/10.1162/tac...

  58. [68]

    Why is ‘Chicago’ deceptive?

    Vivian Lai, Han Liu, and Chenhao Tan. 2020. "Why is ‘Chicago’ deceptive?" Towards Building Model-Driven Tutorials for Humans. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–13...

  59. [69]

    Vivian Lai and Chenhao Tan. 2019. On Human Predictions with Explanations and Predictions of Machine Learning Models: A Case Study on Deception Detection. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). Association for Computing Machin...

  60. [70]

    Placebic

    Ellen Langer, Arthur Black, and Benzio Chanowitz. 1978. The Mindlessness of Ostensibly Thoughtful Action: The Role of "Placebic" Information in Inter- personal Interaction. Journal of Personality and Social Psychology 36, 6 (1978), 635–642. https://doi.org/10.1037/0022-3514.36.6.635

  61. [71]

    Yoonjoo Lee, Kihoon Son, Tae Soo Kim, Jisu Kim, John Joon Young Chung, Eytan Adar, and Juho Kim. 2024. One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations. In Proceedings of the 2024 ACM Conference on Fairness, Accountabilit...

  62. [72]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Gener- ation for Knowledge-Intensive NLP Tasks. In Advances...

  63. [73]

    Q Vera Liao and S Shyam Sundar. 2022. Designing for Responsible Trust in AI Systems: A Communication Perspective. Proceedings of the 2022 Conference on Fairness, Accountability, and Transparency (2022)

  64. [74]

    Q Vera Liao and Kush R Varshney. 2021. Human-Centered Explainable AI (XAI): From Algorithms to User Experiences. arXiv preprint arXiv:2110.10790 (2021)

  65. [75]

    Liquin and Tania Lombrozo

    Emily G. Liquin and Tania Lombrozo. 2022. Motivated to Learn: An Account of Explanatory Satisfaction. Cognitive Psychology 132 (2022), 101453. https: //doi.org/10.1016/j.cogpsych.2021.101453

  66. [76]

    Han Liu, Vivian Lai, and Chenhao Tan. 2021. Understanding the Effect of Out- of-distribution Examples and Interactive Explanations on Human-AI Decision Making. Proc. ACM Hum.-Comput. Interact. 5, CSCW2, Article 408 (oct 2021), 45 pages. https://doi.org/10.1145/3479552

  67. [77]

    Nelson Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating Verifiability in Generative Search Engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singa...

  68. [78]

    Tania Lombrozo. 2006. The Structure and Function of Explanations. Trends in Cognitive Sciences 10, 10 (2024/09/10 2006), 464–470

  69. [79]

    Tania Lombrozo. 2007. Simplicity and Probability in Causal Explanation. Cog- nitive Psychology 55, 3 (2007), 232–257. https://doi.org/10.1016/j.cogpsych.2006. 09.006

  70. [80]

    Tanya Lombrozo. 2012. Explanation and Abductive Inference. In The Oxford Handbook of Thinking and Reasoning . Oxford University Press. https://doi.org/ 10.1093/oxfordhb/9780199734689.013.0014

  71. [81]

    Tania Lombrozo. 2016. Explanatory Preferences Shape Learning and Inference. Trends in Cognitive Sciences 20, 10 (2024/09/10 2016), 748–759

  72. [82]

    Tania Lombrozo and Emily G. Liquin. 2023. Explanation Is Effective Because It Is Selective. Current Directions in Psychological Science 32, 3 (2023), 212–219. https://doi.org/10.1177/09637214231156106

  73. [83]

    Duri Long and Brian Magerko. 2020. What is AI Literacy? Competencies and Design Considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–16. https://doi.org/10.1145/331...

  74. [84]

    Zhuoran Lu and Ming Yin. 2021. Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and Risks. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21) . Association for Computing Machinery, New York, NY, U...

  75. [85]

    Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Mari- anna Apidianaki, and Chris Callison-Burch. 2023. Faithful Chain-of-Thought Reasoning. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of...

  76. [86]

    B. F. Malle and Joshua Knobe. 1997. Which Behaviors Do People Explain? A Basic Actor–Observer Asymmetry. Journal of Personality and Social Psychology 72, 2 (1997), 288

  77. [87]

    Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E. Peters. 2021. Few- Shot Self-Rationalization with Natural Language Prompts. In NAACL-HLT. https://api.semanticscholar.org/CorpusID:244130199

  78. [88]

    Mills, Judith H

    Candice M. Mills, Judith H. Danovitch, Sydney P. Rowles, and Ian L. Campbell

  79. [89]

    Sina Mohseni, Fan Yang, Shiva Pentyala, Mengnan Du, Yi Liu, Nic Lupfer, Xia Hu, Shuiwang Ji, and Eric Ragan. 2021. Machine Learning Explanations to Prevent Overtrust in Fake News Detection. Proceedings of the International AAAI Conference on Web and Social Media 15, 1 (May 202...

  80. [90]

    Hansen Morten Hertzum and Hans H.K

    Kristin D. Hansen Morten Hertzum and Hans H.K. Andersen. 2009. Scrutinis- ing Usability Evaluation: Does Thinking Aloud Affect Behaviour and Men- tal Workload? Behaviour & Information Technology 28, 2 (2009), 165–181. https://doi.org/10.1080/01449290701773842

  81. [91]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  82. [92]

    Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023. The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions. In Proceedings of the 2023 Conference on Empirical Method...

  83. [93]

    Psychonomic Bulletin & Review 24, 5 (2017), 1465–1477

    Children’s Success at Detecting Circular Explanations and Their Interest in Future Learning. Psychonomic Bulletin & Review 24, 5 (2017), 1465–1477

  84. [94]

    Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Yang Wang. 2023. On the Risk of Misinformation Pollution with Large Language Models. In The 2023 Conference on Empirical Methods in Natural Language Processing. https://openreview.net/forum?id=voBhcwDyPt

  85. [95]

    Samir Passi, Shipi Dhanorkar, and Mihaela Vorvoreanu. 2024. Appropriate reliance on Generative AI: Research synthesis . Technical Report MSR-TR-2024-7. Microsoft. https://www.microsoft.com/en-us/research/publication/appropriate- reliance-on-generative-ai-research-synthesis/

  86. [96]

    2022.Overreliance on AI: Literature Review

    Samir Passi and Mihaela Vorvoreanu. 2022.Overreliance on AI: Literature Review. Technical Report MSR-TR-2022-12. Microsoft. https://www.microsoft.com/en- us/research/publication/overreliance-on-ai-literature-review/ CHI ’25, April 26-May 1, 2025, Yokohama, Japan Kim, Vaughan, ...

  87. [97]

    Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wort- man Wortman Vaughan, and Hanna Wallach. 2021. Manipulating and Measuring Model Interpretability. In Proceedings of the 2021 CHI Conference on Human Fac- tors in Computing Systems (CHI ’21). Associatio...

  88. [98]

    Marvin Pafla, Kate Larson, and Mark Hancock. 2024. Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24) . Association f...

  89. [99]

    Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2022. Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges. Statistics Surveys 16, none (2022), 1 – 85. https: //doi.org/10.1214/21-SS133

  90. [100]

    Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Ver- bosity Bias in Preference Labeling by Large Language Models. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following . https://openreview. net/forum?id=magEgFpK1y

  91. [101]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2023. A Missing Piece in the Puzzle: Considering the Role of Task Complexity in Human-AI Decision Making. In Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’23). Association for Compu...

  92. [102]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24)...

  93. [103]

    Danielson

    Soo Young Rieh and David R. Danielson. 2007. Credibility: A Multidisciplinary Framework. Annual Review of Information Science and Technology 41, 1 (2007), 307–364. https://doi.org/10.1002/aris.2007.1440410114

  94. [104]

    Max Schemmer, Patrick Hemmer, Maximilian Nitsche, Niklas Kühl, and Michael Vössing. 2022. A Meta-Analysis of the Utility of Explainable Artificial Intelli- gence in Human-AI Decision-Making. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’22) ....

  95. [105]

    Murray Shanahan. 2024. Talking about Large Language Models. Commun. ACM 67, 2 (Jan 2024), 68–79. https://doi.org/10.1145/3624724

  96. [106]

    Vera Liao, and Ziang Xiao

    Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024. Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information Seeking. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York,...

  97. [107]

    Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023. Large Language Models can be Easily Distracted by Irrelevant Context. In Proceedings of the 40th International Conference on Machine Learning (ICML’23) . JMLR.o...

  98. [108]

    Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Pad- makumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2022. Rome was built in 1776: A Case Study on Factual Correctness in Knowledge-Grounded Response Generation. arXiv:cs.CL/2110.05456 https://arxiv.org/ab...

  99. [109]

    Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, and Jordan Boyd-Graber. 2024. Large Language Models Help Humans Ver- ify Truthfulness – Except When They Are Convincingly Wrong. In Proceed- ings of the 2024 Conference of the North American Chapter ...

  100. [110]

    Rothschild, Daniel G

    Sofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, and Jake M. Hofman. 2023. Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment. arXiv:cs.HC/2307.03744

  101. [111]

    J.D. Trout. 2008. Seduction Without Cause: Uncovering Explanatory Neurophilia. Trends in Cognitive Sciences 12, 8 (2008), 281–282. https://doi.org/10.1016/j.tics. 2008.05.004

  102. [112]

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. Lan- guage Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. InProceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’...

  103. [113]

    Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott We...

  104. [114]

    Bernstein, and Ranjay Krishna

    Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. Proc. ACM Hum.-Comput. Interact. 7, CSCW1, Article 129 (apr 2023), 38 ...

  105. [115]

    Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang

  106. [116]

    Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations . https://open...

  107. [117]

    Xinru Wang and Ming Yin. 2021. Are Explanations Helpful? A Compara- tive Study of the Effects of Explanations in AI-Assisted Decision-Making. In Proceedings of the 26th International Conference on Intelligent User Interfaces (IUI ’21). Association for Computing Machinery, New ...

  108. [118]

    Vera Liao, and Jen- nifer Wortman Vaughan

    Helena Vasconcelos, Gagan Bansal, Adam Fourney, Q. Vera Liao, and Jen- nifer Wortman Vaughan. 2024. Generation Probabilities Are Not Enough: Exploring the Effectiveness of Uncertainty Highlighting in AI-Powered Code Completions. ACM Transactions on Computer-Human Interaction (2024)

  109. [119]

    Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024. How Inter- pretable are Reasoning Explanations from Prompting Large Language Models?. In Findings of the Association for Computational Linguistics: NAACL 2024 , Kevin Duh, Helena Gomez, and Steven Bethard (Eds.)....

  110. [120]

    Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Court- ney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, Willi...

  111. [121]

    Deena Skolnick Weisberg, Jordan C.V Taylor, and Emily J Hopkins. 2015. De- constructing the Seductive Allure of Neuroscience Explanations. Judgment and Decision Making 10, 5 (2015), 429–441

  112. [122]

    Benjamin Weiser and Nate Schweber. 2023. The ChatGPT Lawyer Explains Himself. New York Times (June 2023)

  113. [123]

    Henry M. Wellman. 2011. Reinvigorating Explanations for the Study of Early Cognitive Development. Child Development Perspectives 5, 1 (2011), 33–38. https://doi.org/10.1111/j.1750-8606.2010.00154.x

  114. [124]

    Nadine Wathen and Jacquelyn Burkell

    C. Nadine Wathen and Jacquelyn Burkell. 2002. Believe It or Not: Factors Influ- encing Credibility on the Web. Journal of the American Society for Information Science and Technology 53, 2 (2002), 134–144. https://doi.org/10.1002/asi.10016

  115. [125]

    Jennifer Wortman Vaughan and Hanna Wallach. 2021. A Human-Centered Agenda for Intelligible Machine Learning. In Machines We Trust: Perspectives on Dependable AI, Marcello Pelillo and Teresa Scantamburlo (Eds.). MIT Press

  116. [126]

    Roy Xie, Chengxuan Huang, Junlin Wang, and Bhuwan Dhingra. 2024. Adver- sarial Math Word Problem Generation. arXiv:cs.CL/2402.17916

  117. [127]

    Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023. A Critical Evaluation of Evaluations for Long-form Question Answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Gra...

  118. [128]

    Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the Effect of Accuracy on Trust in Machine Learning Models. In Proceedings of the 2019 ACM CHI Conference on Human Factors in Computing Systems

  119. [129]

    Kun Yu, Shlomo Berkovsky, Ronnie Taib, Jianlong Zhou, and Fang Chen. 2019. Do I Trust My Machine Teammate? An Investigation from Perception to De- cision. In Proceedings of the 24th International Conference on Intelligent User Interfaces (IUI ’19) . Association for Computing M...

  120. [130]

    Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi

  121. [131]

    Zemla, Steven Sloman, Christos Bechlivanidis, and David A

    Jeffrey C. Zemla, Steven Sloman, Christos Bechlivanidis, and David A. Lagnado

  122. [132]

    Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020. Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-assisted Decision Making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 295–305

  123. [133]

    Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for Large Language Models: A Survey. ACM Trans. Intell. Syst. Technol. 15, 2, Article 20 (feb 2024), 38 pages. https://doi.org/10.1145/3639372

  124. [134]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2024. Judging LLM-as-a-Judge with MT- bench and Chatbot Arena. In Proceedings of the 37th Interna...

  125. [135]

    Hwang, Xiang Ren, and Maarten Sap

    Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap. 2024. Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty. arXiv:cs.CL/2401.06730

  126. [136]

    Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji rong Wen. 2023. Large Language Models for Information Retrieval: A Survey. ArXiv abs/2308.07107 (2023). https://api. semanticscholar.org/CorpusID:260887838

  127. [137]

    Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending Against Neural Fake News. In Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R....

  128. [139]

    Psychonomic Bulletin & Review 24, 5 (2017), 1488–1500

    Evaluating Everyday Explanations. Psychonomic Bulletin & Review 24, 5 (2017), 1488–1500. Fostering Appropriate Reliance on Large Language Models CHI ’25, April 26-May 1, 2025, Yokohama, Japan

  129. [145]

    Which animal was sent to space first, cockroach or moon jellyfish?

    Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-Tuning Language Models from Human Preferences. arXiv preprint arXiv:1909.08593 (2019). https: //arxiv.org/abs/1909.08593 CHI ’25, April 26-...

  130. [146]

    https://science.nasa.gov/moon/moon- walkers/ 3

    https://simple.wikipedia.org/wiki/List_of_people_who_have_ walked_on_the_Moon 2. https://science.nasa.gov/moon/moon- walkers/ 3. https://www.discovermagazine.com/planet-earth/what- has-been-found-in-the-deep-waters-of-the-mariana-trench CHI ’25, April 26-May 1, 2025, Yokohama,...

  131. [147]

    https://www.healthline.com/health/hair-density • Incorrect: Yes, gorillas have twice as many hairs per square inch as humans

    https://www.nationalgeographic.com/science/article/the-semi- naked-ape-or-why-peach-fuzz-makes-it-harder-for-parasites 3. https://www.healthline.com/health/hair-density • Incorrect: Yes, gorillas have twice as many hairs per square inch as humans. Gorillas have a significantly...

  132. [148]

    https://midwesteyecenter

    https://2020visioncare.com/the-eye-a-marvel-of-complexity- with-over-2-million-working-parts/ 2. https://midwesteyecenter. com/what-are-the-makings-of-the-human-eye/ 3. https://www. optometrists.org/general-practice-optometry/guide-to-eye-health/ how-does-the-eye-work/ • Incor...

  133. [149]

    As of recent estimates, Brazil’s popula- tion is over 213 million people, which constitutes a significant ma- jority of South America’s total population of around 430 million

    https://www.worldometers.info/world-population/south-america- population/ • Incorrect: Yes, more than two-thirds of South America’s popula- tion live in Brazil because Brazil is the largest and most populous country on the continent. As of recent estimates, Brazil’s popula- ti...

  134. [2008]

    In Proceedings of ACL-08: HLT , Johanna D

    Finding Contradictions in Text. In Proceedings of ACL-08: HLT , Johanna D. Moore, Simone Teufel, James Allan, and Sadaoki Furui (Eds.). Association for Computational Linguistics, Columbus, Ohio, 1039–1047. https://aclanthology. org/P08-1118

  135. [2017]

    Psychonomic Bulletin & Review 24, 5 (2017), 1451–1464

    Concreteness and Abstraction in Everyday Explanation. Psychonomic Bulletin & Review 24, 5 (2017), 1451–1464. https://doi.org/10.3758/s13423-017- 1299-3

  136. [2019]

    In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (CHI EA ’19)

    The Impact of Placebic Explanations on Trust in Intelligent Systems. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (CHI EA ’19) . Association for Computing Machinery, New York, NY, USA, 1–6. https://doi.org/10.1145/3290607.3312787

  137. [2022]

    Reframing Human-AI Collaboration for Generating Free-Text Explana- tions. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Marine Carpuat, Marie-Catherine de Marneffe, and Ivan V...

  138. [2023]

    arXiv:cs.CL/2310.07521

    Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity. arXiv:cs.CL/2310.07521

  139. [2024]

    arXiv:cs.CL/2311.08596 https://arxiv.org/abs/2311.08596

    Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment. arXiv:cs.CL/2311.08596 https://arxiv.org/abs/2311.08596

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.