Pith. sign in

REVIEW 5 major objections 5 minor 1 cited by

Engaging with AI: How Interface Design Shapes Human-AI Collaboration in High-Stakes Decision-Making

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Across six interface mechanisms tested with 108 participants, confidence levels and textual explanations improved human-AI decision accuracy, while reflective mechanisms such as AI-driven questions reduced it.

desk verdict A useful empirical comparison undermined by an unsourced ground-truth rule, internal contradictions, and a confidence-level leak; reject but invite major revision. read the letter →

arxiv 2501.16627 v1 pith:C3WHTC6C submitted 2025-01-28 cs.HC cs.AI

classification cs.HCcs.AI
keywords human-AIcollaborationexplainableAIcognitiveforcingfunctionstrustcalibrationdecisionaccuracyloaddiabetesmanagementinterfacedesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a controlled experiment on how the design of AI decision-support interfaces changes outcomes when people and an AI choose meals for a Type 2 diabetes patient. It claims that the mechanisms that improved decision accuracy were the transparent, low-effort ones—AI confidence levels and textual explanations—while mechanisms that forced deeper reflection, such as human feedback and AI-driven questions, reduced accuracy by raising cognitive effort. Performance visualization improved accuracy only at borderline significance and had mixed effects on trust. The paper explains these differences with a quantity called Explanation Information Load (EIL), a logarithmic count of the elements an interface asks a user to process, paired with whether the mechanism supplies interpretable model reasoning. If the claim is right, interface designers in similar high-stakes settings should lead with confidence levels and concise explanations and treat reflective interventions as a cognitive tax.

What carries the argument

The central analytical object is Explanation Information Load (EIL), defined as the sum, over the distinct information-carrying components of an explanation, of $\log_2$ of the number or length of those components, with the image-information term removed when comparing conditions. The paper reports EIL values from 0.602 for confidence levels and performance visualization to 1.965 for AI-driven questions. EIL operationalizes cognitive load as a count-based information-theoretic quantity, and it is paired with a qualitative condition: whether the mechanism provides interpretable model reasoning, such as a text rationale or confidence numbers, rather than raw information. Together these two components carry the argument: low-to-moderate EIL with interpretable reasoning predicts accuracy gains, while high EIL without interpretive support does not. The design and interpretation also rely on dual-process theory, with mechanisms intended to shift users from fast System 1 thinking to more deliberate System 2 evaluation.

What would settle it

Re-run the six condition comparisons using a panel of certified diabetes nutritionists' labels for the same 20 meal pairs instead of the paper's threshold rule; if the ranking of mechanisms changes, the accuracy conclusions are an artifact of the hand-built ground truth. Independently, recompute the Table 2 EIL values from the printed formulas using the actual interface elements, since the reported numbers do not visibly match those formulas.

Watch

Extended reading notes

Core claim

The central claim is that six decision-support mechanisms fall into a clear ordering on human-AI task performance, and that this ordering is explained by the interaction of Explanation Information Load (EIL) with the presence of interpretable model reasoning. In a 20-question, three-phase meal-selection experiment with 18 participants per condition, AI confidence levels significantly raised decision accuracy (p = 0.004) and visual explanations did so as well (p = 0.018), while textual explanations improved engagement and perceived reliability. Human feedback (p = 0.674) and AI-driven questions (p = 0.302) did not improve accuracy, and AI-driven questions significantly reduced trust (p = 0.032). The paper therefore argues that mechanisms with manageable EIL and interpretable outputs—confidence levels, textual explanations, and to a lesser extent performance visualization—achieve the best balance between fostering engagement and maintaining performance. A striking additional finding is that none of the six mechanisms increased trust overall, which the paper reads as users treating the explanations as learning tools rather than as reasons to rely more on the AI.

Load-bearing premise

The accuracy ranking rests on a hand-assembled ground-truth rule—a meal counts as correct for blood-sugar control only when protein is 20–32% of calories and fat 20–35%, with lower carbohydrate as the tie-breaker—and on the reported cognitive-load values following from the printed formulas; if either link fails, the conclusion that confidence levels and text explanations improve accuracy has no foundation.

Editorial extensions

If this is right

  • Interface designers for AI-assisted decisions should prefer low-EIL, interpretable mechanisms such as AI confidence levels and text explanations when the goal is decision accuracy and trust calibration.
  • Reflective interventions like AI-driven questions and human feedback should be deployed only when the task's stakes justify the extra cognitive effort, since in this study they lowered accuracy and sometimes trust.
  • Visual explanations alone are unlikely to shift trust; they need to be paired with interpretable reasoning cues or other engagement mechanisms.
  • EIL can serve as an early screening metric for proposed explanation mechanisms: manageable load plus interpretable reasoning is the combination to aim for.
  • None of the six mechanisms increased trust, so designers should not expect transparency features to make users rely more on the AI; they may instead strengthen the user's own judgment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the confidence-level result generalizes, it suggests a cheap design heuristic: expose the AI's uncertainty before asking the user to commit, rather than adding longer explanations.
  • The ranking should be re-run against clinician-labeled ground truth for the same meal pairs; if the ranking survives, the conclusion is about real meal quality rather than the paper's specific thresholds.
  • The benefit of confidence levels probably depends on the AI's confidence being honest; with systematically overconfident AI, the mechanism could amplify automation bias instead of correcting it.
  • Because no condition increased trust, a natural longitudinal test is whether users who saw confidence levels or text explanations make better independent decisions later, when the AI is removed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper reports a between-subjects experiment (N=108, 18 per condition) comparing six decision-support mechanisms in a diabetes meal-choice task: text explanations, visual explanations, AI confidence levels, human feedback, AI-driven questions, and performance visualization. Each participant completed three phases (independent choice, AI suggestion without explanation, AI suggestion plus a mechanism), and the authors measure decision accuracy against a ground-truth rule, self-reported trust, satisfaction, reliability, complexity, and engagement. The paper claims that text explanations, AI confidence levels, and performance visualization improved human-AI collaborative performance, that human feedback and AI-driven questions reduced performance, and that a balanced Explanation Information Load (EIL) with interpretable outputs is the best design principle. The central statistical evidence for the accuracy-based claims depends on a hand-built ground-truth rule (Section 6.9.2) that is not directly sourced, and several sections report directly contradictory results for the same conditions.

Significance. If the results were valid, the study would usefully inform the design of XAI and cognitive forcing functions in high-stakes decision support, and the controlled six-condition design with 108 participants is a reasonable empirical contribution. The paper also has strengths: it uses objective and subjective measures, reports non-parametric tests, includes a power analysis, and states that data and code will be released. However, the current accuracy metric is based on an unvalidated and clinically questionable ground-truth rule, the EIL values are not reproducible from the stated formulas, and multiple sections contradict one another on which mechanisms improved performance. These problems are load-bearing because the abstract and conclusion draw design recommendations directly from the affected comparisons.

major comments (5)
  1. [Section 6.9.2 (Eqs. 10–15)] The correctness label used for every accuracy statistic is defined by an unvalidated rule: 20% ≤ PP ≤ 32% and 20% ≤ FP ≤ 35%, with lower carbohydrate percentage used only as a tie-breaker. The text cites ADA references [113,114], but those references do not provide these exact threshold values, and the rule ignores carbohydrate quantity and quality unless both meals are already 'balanced', even though carbohydrates are the primary determinant of postprandial glucose. Because P1/P2/P3 accuracy, the McNemar tests, the Wilcoxon tests, and all between-condition performance comparisons are computed against this binary label, an arbitrary or incorrect ground truth would reorder the paper's conclusions. The authors must provide a direct source or derivation for the thresholds, validate the rule against clinical guidance or dietary indices, and report a sensitivity analysis under plausible alternative thresholds.
  2. [Sections 2.3, 2.5, 3.1.1] The performance results are internally contradictory. Section 2.3 reports that visual explanations (C2) significantly increase user accuracy (p = 0.018) and assigns C2 a moderate EIL of 1.749; two paragraphs later it states that C2 (EIL = 1.520) shows no significant improvement (p = 0.673), and Section 3.1.1 states that C2 does not significantly improve user performance. Similarly, Section 2.3 reports no significant improvement for textual explanations C1 (p = 0.673), while Sections 3.1.1 and 3.3.1 claim that C1 significantly improves user performance. No table in the paper reports the accuracy p-values used in Section 2.3, so the reader cannot determine which statement reflects the actual data. These contradictions concern the paper's central claim about which mechanisms improve human-AI task performance.
  3. [Section 6.8, Table 2] The EIL values in Table 2 are not reproducible from Eqs. (1)–(6). For example, Eq. (3) gives log2(NCL); with two confidence levels this equals 1.0, but the table reports 0.602 for C3. Eq. (6) with four chart components gives 2.0 under log2, while 0.602 equals log10(4). No base, scaling, or additional normalization is specified that reconciles all six entries with the stated formulas. Since hypothesis H3 is evaluated by comparing EIL values across conditions, the metric must be defined precisely and the table must be recomputed from the formulas.
  4. [Section 2.2, Table 6] Table 6 compares 'engagement changes' between P1–P2 and P1–P3. For the significant rows (C1, C3, C4, C6), the P3 change is smaller than the P2 change (e.g., C3: 0.183 vs. 0.033; C6: 0.456 vs. 0.394). The text interprets these p-values as evidence that the mechanism increases engagement, but if these are paired differences in the accuracy-change measure, the significant comparison indicates that the mechanism attenuates the improvement seen in P2, not that it increases engagement. The measure itself is also undefined in the text: it is called 'engagement change' but the values look like accuracy deltas. This ambiguity undermines the support for H2 and requires clarification or reanalysis.
  5. [Sections 2.8, 2.9, Table 3] Reported within-condition p-values are inconsistent with Table 3. Section 2.8 states that AI-driven questions (C5) significantly decrease trust (p = 0.032), but Table 3 reports p = 0.100 for C5 trust; Section 2.9 states that performance visualization (C6) shows a non-significant trust change (p = 0.099), while Table 3 reports p = 0.032 for C6 trust. A similar swap occurs for system complexity: Section 2.9 says C6 is non-significant (p = 0.218), while Table 3 reports p = 0.008, and Section 2.8 reports C5 as significant at p = 0.008 when Table 3 shows p = 0.218. These errors affect the trust and complexity conclusions in the Discussion and must be corrected.
minor comments (5)
  1. [Section 2.5] The section heading 'The Impact of Text Explanations on Decision-Making' appears to be a typo; the section describes visual explanations (C2) and should be retitled accordingly.
  2. [Section 6.8.2, Eq. (2)] Equation (2) has a malformed expression: the sentence 'The EIL for visual explanations is calculated as follows' is repeated, and the summation has mismatched parentheses that make the intended formula unclear.
  3. [Table 6 caption] The caption says the table compares 'engagement changes' but the values appear to be accuracy changes; the paper should define the variable being compared and explain why it is labeled engagement.
  4. [Section 2.3] The accuracy p-values cited in Section 2.3 (e.g., p = 0.018, p = 0.004, p = 0.673) do not appear in any results table; the authors should include the full per-condition accuracy statistics so the reported effects are verifiable.
  5. [Section 2.8] The phrase 'satisfaction slightly decreases... (p = 0.773)' in the AI-driven questions condition is followed by text that appears to discuss a different condition; the paragraph should be checked for consistency with Table 3.

Circularity Check

1 steps flagged · score 6.0 of 10

AI-confidence result is partially circular: the "user CL" shown to participants is generated from the same ground-truth label used to compute decision accuracy.

  1. fitted input called prediction [Section 6.3.1 (Prompt for Calculating User's CL); Algorithm 1 in Section 6.9.1; result reported in Sections 2.3 and 2.6 (Condition C3)]
    "New Query for Prediction: ... The ground truth is: {new_ground_truth}. Instruction: Return the predicted likelihood as a percentage (0-100) of the user choosing the correct meal... Algorithm 1: Determine ground truth GTnew for current query; Append current query information to Pdata: 'New Query: M1, M2, Ground Truth = GTnew.'"

    The user confidence level displayed in Condition C3 is called a prediction, but the prompt that generates it is given the current query's ground truth (GTnew). The same ground truth is the label used to compute every decision-accuracy score (Equations 10-15, Section 6.9.2), including the significant C3 accuracy improvement (p = 0.004) that supports the claim that AI confidence levels enhance performance. The displayed user CL is therefore not an independent decision aid; it is a function of the outcome label itself. Any performance benefit it produces is forced by construction, because the intervention contains information about the answer key, so the comparison does not cleanly test the mechanism.

full rationale

Most of this paper is a conventional between-subjects user study and is not circular: the six interface mechanisms are not fitted to the accuracy outcomes, and the only same-author citation ([50], an XplainLLM dataset) is an incidental reference for text interpretability rather than a load-bearing premise. The ground-truth rule (Equations 10-15) is under-sourced and clinically questionable, but that is a validity and measurement concern rather than a circular derivation, since it is an external input assumption. The EIL values in Table 2 also do not reproduce the printed formulas (e.g., Equation 3 with NCL = 2 gives log2(2) = 1, not 0.602; Equation 6 with Nchart = 4 gives log2(4) = 2, not 0.602), which is a reproducibility problem, but the paper uses EIL as an explanatory post-hoc construct rather than deriving performance from it, so I do not count that as a demonstrated circular reduction. The genuine circular step is the user-CL prediction in Condition C3: Algorithm 1 and the prompt in Section 6.3.1 include the current query's ground truth in the input to the model that produces the user confidence level shown to participants. Since the same ground truth defines the accuracy outcome, the C3 performance gain is not independent evidence for the mechanism; the prediction reduces to its label input by construction. This makes the AI-confidence-level result partially circular while leaving the text-explanation, performance-visualization, and other condition claims largely independent.

Assumptions & free parameters 3 free parameters · 5 assumptions · 1 invented entities

The paper's central claims rest on author-chosen ground-truth thresholds, an unvalidated cognitive-load metric, and the assumption that GPT-4o verbalized confidence is meaningful. The EIL metric is introduced for this paper and is not externally validated.

free parameters (3)
  • Protein percentage balance thresholds = 20% to 32%
    Protein percentage range for ground-truth 'balanced' meal; asserted in Equations 10-11 without direct citation of a source specifying these exact values.
  • Fat percentage balance thresholds = 20% to 35%
    Fat percentage range used in the binary balance indicator B, Equations 10-11; author-chosen thresholds that define the correct answer.
  • Power analysis effect size = r = 0.60
    Assumed effect size for sample size calculation; not estimated from pilot data.
assumptions (5)
  • domain assumption Dual process theory (System 1/System 2) correctly describes user engagement with AI explanations.
    The entire framework and mechanism design is premised on this psychological theory (Section 1).
  • ad hoc to paper Information load is proportional to the base-2 (or base-10) logarithm of the number of interface elements.
    EIL formulas in Section 6.8 assume a count-based information-theoretic measure of cognitive load; no validation is provided.
  • ad hoc to paper Accuracy change from P1 to P3 compared with P1 to P2 is a valid measure of engagement.
    The paper uses accuracy deltas as 'engagement' (Table 6, Section 2.2), though engagement is a cognitive state, not an outcome.
  • domain assumption The binary balance indicator B and lower-carbohydrate tie-break correctly identify the diabetes-appropriate meal.
    Ground truth is defined by author-selected macronutrient thresholds (Equations 10-15); if these do not reflect real dietary guidance, all accuracy scores are invalid.
  • domain assumption GPT-4o's stated confidence levels correspond to meaningful AI uncertainty.
    The AI's CL in C3 is the model's verbalized confidence, not a calibrated probability; no calibration analysis is reported.
invented entities (1)
  • Explanation Information Load (EIL)
    purpose: Quantify the cognitive processing load of each explanation mechanism (Section 6.8).
    EIL is a new composite metric built from element counts; it is never validated against cognitive-load questionnaires or behavioral measures, and its formulas are inconsistent (log2 vs log10).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Engaging with AI: How Interface Design Shapes Human-AI Collaboration in High-Stakes Decision-Making." pith.science (2026). https://pith.science/paper/C3WHTC6C

@misc{pith2026250116627,
  author       = {Pith},
  title        = {Pith review of: Engaging with AI: How Interface Design Shapes Human-AI Collaboration in High-Stakes Decision-Making},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C3WHTC6C}},
  note         = {Machine review of arXiv:2501.16627}
}
read the original abstract

As reliance on AI systems for decision-making grows, it becomes critical to ensure that human users can appropriately balance trust in AI suggestions with their own judgment, especially in high-stakes domains like healthcare. However, human + AI teams have been shown to perform worse than AI alone, with evidence indicating automation bias as the reason for poorer performance, particularly because humans tend to follow AI's recommendations even when they are incorrect. In many existing human + AI systems, decision-making support is typically provided in the form of text explanations (XAI) to help users understand the AI's reasoning. Since human decision-making often relies on System 1 thinking, users may ignore or insufficiently engage with the explanations, leading to poor decision-making. Previous research suggests that there is a need for new approaches that encourage users to engage with the explanations and one proposed method is the use of cognitive forcing functions (CFFs). In this work, we examine how various decision-support mechanisms impact user engagement, trust, and human-AI collaborative task performance in a diabetes management decision-making scenario. In a controlled experiment with 108 participants, we evaluated the effects of six decision-support mechanisms split into two categories of explanations (text, visual) and four CFFs. Our findings reveal that mechanisms like AI confidence levels, text explanations, and performance visualizations enhanced human-AI collaborative task performance, and improved trust when AI reasoning clues were provided. Mechanisms like human feedback and AI-driven questions encouraged deeper reflection but often reduced task performance by increasing cognitive effort, which in turn affected trust. Simple mechanisms like visual explanations had little effect on trust, highlighting the importance of striking a balance in CFF and XAI design.

Figures

Figures reproduced from arXiv: 2501.16627 by the authors.

Figure 1
Figure 1. Dual-layer framework guiding the transition from System 1 (intuitive thinking) to System 2 (deliberative [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Visualization of User Answer Correctness. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Questionnaire scores across six conditions on the five measures: satisfaction, system complexity, reliability, [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Interface design across different phases: Phase 1, where the user makes an independent decision without AI [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: Pre- and post-assessment ratings across different criteria: Satisfaction, Complexity, Reliability, Trust, and [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 6
Figure 6. Figure 6: Scatter plots illustrating the trust scores of participants across six conditions in pre- and post-questionnaire. [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation

    cs.AI 2026-08 conditional novelty 5.0 of 10

    A mixed-stakeholder UK workshop rated 13 AI policing use cases, rejecting recidivism risk assessment outright while accepting most others conditionally, and found that a racial-equity focus broadened, not narrowed, th...

Reference graph

Works this paper leans on

115 extracted references · 68 canonical work pages · cited by 1 Pith paper

  1. [1]

    Elements of chronic disease management service system: an empirical study from large hospitals in china

    Shuzhen Zhao, Renjie Du, Yanhua He, Xiaoli He, Yaxin Jiang, and Xinli Zhang. Elements of chronic disease management service system: an empirical study from large hospitals in china. Scientific reports, 12(1):5693, 2022

  2. [2]

    Characteristics of self-care inter- ventions for patients with a chronic condition: A scoping review

    Barbara Riegel, Heleen Westland, Paolo Iovino, Ingrid Barelds, Joyce Bruins Slot, Michael A Stawnychy, Onome Osokpo, Elise Tarbi, Jaap CA Trappenburg, Ercole Vellone, et al. Characteristics of self-care inter- ventions for patients with a chronic condition: A scoping review. International journal of nursing studies , 116:103713, 2021

  3. [3]

    Artificial intelligence: the future for diabetes care

    Samer Ellahham. Artificial intelligence: the future for diabetes care. The American journal of medicine , 133(8):895–900, 2020

  4. [4]

    Type 2 diabetes and cognitive dysfunction—towards effective management of both comorbidities

    Velandai Srikanth, Alan J Sinclair, Felicia Hill-Briggs, Chris Moran, and Geert Jan Biessels. Type 2 diabetes and cognitive dysfunction—towards effective management of both comorbidities. The lancet Diabetes & en- docrinology, 8(6):535–545, 2020

  5. [5]

    Self-management: a comprehensive approach to management of chronic conditions

    Patricia A Grady and Lisa Lucio Gough. Self-management: a comprehensive approach to management of chronic conditions. American journal of public health, 104(8):e25–e31, 2014

  6. [6]

    From Glucose Patterns to Health Outcomes: A Generalizable Foundation Model for Continuous Glucose Monitor Data Analysis

    Guy Lutsker, Gal Sapir, Anastasia Godneva, Smadar Shilo, Jerry R Greenfield, Dorit Samocha-Bonet, Shie Mannor, Eli Meirom, Gal Chechik, Hagai Rossman, et al. From glucose patterns to health outcomes: A gen- eralizable foundation model for continuous glucose monitor data analysis. arXiv preprint arXiv:2408.11876, 2024

  7. [7]

    Data-driven allocation of preventive care with application to diabetes mellitus type ii

    Mathias Kraus, Stefan Feuerriegel, and Maytal Saar-Tsechansky. Data-driven allocation of preventive care with application to diabetes mellitus type ii. Manufacturing & Service Operations Management , 26(1):137–153, 2024

  8. [8]

    Ai-supported insulin dosing for type 2 diabetes

    Georgia M Davis, Hui Shao, and Francisco J Pasquel. Ai-supported insulin dosing for type 2 diabetes. Nature Medicine, 29(10):2414–2415, 2023. 30 Engaging with AI

Show all 115 references
  1. [9]

    Machine learning tools for long-term type 2 diabetes risk prediction

    Nikos Fazakis, Otilia Kocsis, Elias Dritsas, Sotiris Alexiou, Nikos Fakotakis, and Konstantinos Moustakas. Machine learning tools for long-term type 2 diabetes risk prediction. ieee Access, 9:103737–103757, 2021

  2. [10]

    Integrated image-based deep learning and lan- guage models for primary diabetes care

    Jiajia Li, Zhouyu Guan, Jing Wang, Carol Y Cheung, Yingfeng Zheng, Lee-Ling Lim, Cynthia Ciwei Lim, Paisan Ruamviboonsuk, Rajiv Raman, Leonor Corsino, et al. Integrated image-based deep learning and lan- guage models for primary diabetes care. Nature medicine, pages 1–11, 2024

  3. [11]

    Open (clinical) llms are sensitive to instruction phrasings

    Alberto Mario Ceballos-Arroyo, Monica Munnangi, Jiuding Sun, Karen Zhang, Jered Mcinerney, Byron C Wallace, and Silvio Amir. Open (clinical) llms are sensitive to instruction phrasings. InProceedings of the 23rd Workshop on Biomedical Natural Language Processing, pages 50–71, 2024

  4. [12]

    Recommenda- tions for initial diabetic retinopathy screening of diabetic patients using large language model-based artificial intelligence in real-life case scenarios

    Nikhil Gopalakrishnan, Aishwarya Joshi, Jay Chhablani, Naresh Kumar Yadav, Nikitha Gurram Reddy, Pad- maja Kumari Rani, Ram Snehith Pulipaka, Rohit Shetty, Shivani Sinha, Vishma Prabhu, et al. Recommenda- tions for initial diabetic retinopathy screening of diabetic patients us...

  5. [13]

    Llm-powered multimodal ai conversations for diabetes prevention

    Dung Dao, Jun Yi Claire Teo, Wenru Wang, and Hoang D Nguyen. Llm-powered multimodal ai conversations for diabetes prevention. In Proceedings of the 1st ACM Workshop on AI-Powered Q&A Systems for Multimedia, pages 1–6, 2024

  6. [14]

    Multimodal llms for health grounded in individual- specific data

    Anastasiya Belyaeva, Justin Cosentino, Farhad Hormozdiari, Krish Eswaran, Shravya Shetty, Greg Corrado, Andrew Carroll, Cory Y McLean, and Nicholas A Furlotte. Multimodal llms for health grounded in individual- specific data. In Workshop on Machine Learning for Multimodal Heal...

  7. [15]

    A trust based framework for the envelopment of medical ai

    Lena Christine Zuchowski, Matthias Lukas Zuchowski, and Eckhard Nagel. A trust based framework for the envelopment of medical ai. npj Digital Medicine, 7(1):230, 2024

  8. [16]

    Trust in ai and its role in the acceptance of ai technologies

    Hyesun Choung, Prabu David, and Arun Ross. Trust in ai and its role in the acceptance of ai technologies. International Journal of Human–Computer Interaction, 39(9):1727–1739, 2023

  9. [17]

    Trust and medical ai: the challenges we face and the expertise needed to overcome them

    Thomas P Quinn, Manisha Senadeera, Stephan Jacobs, Simon Coghlan, and Vuong Le. Trust and medical ai: the challenges we face and the expertise needed to overcome them. Journal of the American Medical Informatics Association, 28(4):890–894, 2021

  10. [18]

    People over trust ai-generated medical responses and view them to be as valid as doctors, despite low accuracy

    Shruthi Shekar, Pat Pataranutaporn, Chethan Sarabu, Guillermo A Cecchi, and Pattie Maes. People over trust ai-generated medical responses and view them to be as valid as doctors, despite low accuracy. arXiv preprint arXiv:2408.15266, 2024

  11. [19]

    Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making

    Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making. In Proceedings of the 2020 conference on fairness, account- ability, and transparency, pages 295–305, 2020

  12. [20]

    Human reliance on machine learning models when performance feedback is limited: Heuristics and risks

    Zhuoran Lu and Ming Yin. Human reliance on machine learning models when performance feedback is limited: Heuristics and risks. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , pages 1–16, 2021

  13. [21]

    Artificial intelligence trust, risk and security management (ai trism): Frameworks, applications, challenges and future research directions

    Adib Habbal, Mohamed Khalif Ali, and Mustafa Ali Abuzaraida. Artificial intelligence trust, risk and security management (ai trism): Frameworks, applications, challenges and future research directions. Expert Systems with Applications, 240:122442, 2024

  14. [22]

    How transparency modulates trust in artificial intelligence

    John Zerilli, Umang Bhatt, and Adrian Weller. How transparency modulates trust in artificial intelligence. Patterns, 3(4), 2022

  15. [23]

    Bridging the gap between ethics and practice: guidelines for reliable, safe, and trustworthy human-centered ai systems

    Ben Shneiderman. Bridging the gap between ethics and practice: guidelines for reliable, safe, and trustworthy human-centered ai systems. ACM Transactions on Interactive Intelligent Systems (TiiS), 10(4):1–31, 2020

  16. [24]

    Designing interpretable ml system to enhance trust in healthcare: A systematic review to proposed responsible clinician-ai-collaboration framework

    Elham Nasarian, Roohallah Alizadehsani, U Rajendra Acharya, and Kwok-Leung Tsui. Designing interpretable ml system to enhance trust in healthcare: A systematic review to proposed responsible clinician-ai-collaboration framework. Information Fusion, page 102412, 2024

  17. [25]

    Artificial intelligence and human trust in healthcare: focus on clinicians

    Onur Asan, Alparslan Emrah Bayrak, Avishek Choudhury, et al. Artificial intelligence and human trust in healthcare: focus on clinicians. Journal of medical Internet research, 22(6):e15154, 2020

  18. [26]

    Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision-making

    Shuai Ma, Ying Lei, Xinru Wang, Chengbo Zheng, Chuhan Shi, Ming Yin, and Xiaojuan Ma. Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision-making. In Proceedings of the 2023 CHI Conference on Huma...

  19. [27]

    To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proceedings of the ACM on Human-computer Interaction, 5(CSCW1):1–21, 2021. 31 Engaging with AI

  20. [28]

    Does the whole exceed its parts? the effect of ai explanations on complementary team performance

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in...

  21. [29]

    Overconfident and unconfident ai hinder human-ai collaboration

    Jingshu Li, Yitian Yang, and Yi-chieh Lee. Overconfident and unconfident ai hinder human-ai collaboration. arXiv preprint arXiv:2402.07632, 2024

  22. [30]

    Do people engage cognitively with ai? impact of ai assistance on incidental learning

    Krzysztof Z Gajos and Lena Mamykina. Do people engage cognitively with ai? impact of ai assistance on incidental learning. In Proceedings of the 27th International Conference on Intelligent User Interfaces , pages 794–806, 2022

  23. [31]

    Explainable arti- ficial intelligence (xai): What we know and what is left to attain trustworthy artificial intelligence

    Sajid Ali, Tamer Abuhmed, Shaker El-Sappagh, Khan Muhammad, Jose M Alonso-Moral, Roberto Con- falonieri, Riccardo Guidotti, Javier Del Ser, Natalia Díaz-Rodríguez, and Francisco Herrera. Explainable arti- ficial intelligence (xai): What we know and what is left to attain trust...

  24. [32]

    Explainable ai (xai): Core ideas, techniques, and solutions.ACM Computing Surveys, 55(9):1–33, 2023

    Rudresh Dwivedi, Devam Dave, Het Naik, Smiti Singhal, Rana Omer, Pankesh Patel, Bin Qian, Zhenyu Wen, Tejal Shah, Graham Morgan, et al. Explainable ai (xai): Core ideas, techniques, and solutions.ACM Computing Surveys, 55(9):1–33, 2023

  25. [33]

    Dissenting explanations: Leveraging disagreement to reduce model overreliance

    Omer Reingold, Judy Hanwen Shen, and Aditi Talati. Dissenting explanations: Leveraging disagreement to reduce model overreliance. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 21537–21544, 2024

  26. [34]

    Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, and Mark O Riedl

    Upol Ehsan, Samir Passi, Q. Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, and Mark O Riedl. The who in xai: How ai background shapes perceptions of ai explanations. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , CHI ’24, New York, NY ,...

  27. [35]

    Understanding user reliance on ai in assisted decision-making

    Shiye Cao and Chien-Ming Huang. Understanding user reliance on ai in assisted decision-making. Proc. ACM Hum.-Comput. Interact., 6(CSCW2), November 2022

  28. [36]

    Algorithm appreciation: People prefer algorithmic to human judgment

    Jennifer M Logg, Julia A Minson, and Don A Moore. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151:90–103, 2019

  29. [37]

    Thinking, fast and slow

    Daniel Kahneman. Thinking, fast and slow. Farrar, Straus and Giroux, 2011

  30. [38]

    Optimizing human-ai collaboration: Effects of motivation and accuracy information in ai-supported decision-making

    Simon Eisbach, Markus Langer, and Guido Hertel. Optimizing human-ai collaboration: Effects of motivation and accuracy information in ai-supported decision-making. Computers in Human Behavior: Artificial Humans, 1(2):100015, 2023

  31. [39]

    Designing for responsible trust in ai systems: A communication perspective

    Q Vera Liao and S Shyam Sundar. Designing for responsible trust in ai systems: A communication perspective. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1257–1268, 2022

  32. [40]

    Human-centered explainable ai (xai): From algorithms to user experiences

    Q Vera Liao and Kush R Varshney. Human-centered explainable ai (xai): From algorithms to user experiences. arXiv preprint arXiv:2110.10790, 2021

  33. [41]

    Explanations can reduce overreliance on ai systems during decision-making

    Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S Bern- stein, and Ranjay Krishna. Explanations can reduce overreliance on ai systems during decision-making. Pro- ceedings of the ACM on Human-Computer Interaction, 7(CSCW1):1–38, 2023

  34. [42]

    Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal

    Valerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal. Understanding the role of human intuition on reliance in human-ai decision-making with explanations. Proc. ACM Hum.-Comput. Interact. , 7(CSCW2), October 2023

  35. [43]

    Dual-process cognitive interven- tions to enhance diagnostic reasoning: a systematic review

    Kathryn Ann Lambe, Gary O’Reilly, Brendan D Kelly, and Sarah Curristan. Dual-process cognitive interven- tions to enhance diagnostic reasoning: a systematic review. BMJ quality & safety, 25(10):808–820, 2016

  36. [44]

    The role of explanations on trust and reliance in clinical decision support systems

    Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. The role of explanations on trust and reliance in clinical decision support systems. In 2015 international conference on healthcare informatics, pages 160–169. IEEE, 2015

  37. [45]

    How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection

    Maia Jacobs, Melanie F Pradier, Thomas H McCoy Jr, Roy H Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry, 11(1):108, 2021

  38. [46]

    On human predictions with explanations and predictions of machine learning models: A case study on deception detection

    Vivian Lai and Chenhao Tan. On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In Proceedings of the conference on fairness, accountability, and transparency, pages 29–38, 2019. 32 Engaging with AI

  39. [47]

    Wilcoxon signed-rank test

    Robert F Woolson. Wilcoxon signed-rank test. Encyclopedia of Biostatistics, 8, 2005

  40. [48]

    Dealing with information overload: a comprehen- sive review

    Miriam Arnold, Mascha Goldschmitt, and Thomas Rigotti. Dealing with information overload: a comprehen- sive review. Frontiers in psychology, 14:1122200, 2023

  41. [49]

    Explainable artificial intelligence for mental health through transparency and interpretability for understandability

    Dan W Joyce, Andrey Kormilitzin, Katharine A Smith, and Andrea Cipriani. Explainable artificial intelligence for mental health through transparency and interpretability for understandability. npj Digital Medicine, 6(1):6, 2023

  42. [50]

    Xplainllm: A qa explanation dataset for understanding llm decision-making

    Zichen Chen, Jianda Chen, Mitali Gaidhani, Ambuj Singh, and Misha Sra. Xplainllm: A qa explanation dataset for understanding llm decision-making. arXiv preprint arXiv:2311.08614, 2023

  43. [51]

    Sunnie S. Y . Kim, Nicole Meister, Vikram V . Ramaswamy, Ruth Fong, and Olga Russakovsky. HIVE: Evalu- ating the human interpretability of visual explanations. In European Conference on Computer Vision (ECCV), 2022

  44. [52]

    Effects of multimodal explanations for autonomous driving on driving performance, cognitive load, expertise, confidence, and trust

    Robert Kaufman, Jean Costa, and Everlyne Kimani. Effects of multimodal explanations for autonomous driving on driving performance, cognitive load, expertise, confidence, and trust. Scientific Reports, 14, 2024

  45. [53]

    Dynamic explanation selection towards successful user-decision support with explainable ai

    Yosuke Fukuchi and Seiji Yamada. Dynamic explanation selection towards successful user-decision support with explainable ai. arXiv preprint arXiv:2402.18016, 2024

  46. [54]

    Incremental xai: Memorable understanding of ai with incremental explanations

    Jessica Y Bo, Pan Hao, and Brian Y Lim. Incremental xai: Memorable understanding of ai with incremental explanations. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1–17, 2024

  47. [55]

    Towards human-centered explainable ai: A survey of user studies for model explanations

    Yao Rong, Tobias Leemann, Thai-Trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav Unhelkar, Tina Seidel, Gjergji Kasneci, and Enkelejda Kasneci. Towards human-centered explainable ai: A survey of user studies for model explanations. IEEE transactions on pattern analysis and mach...

  48. [56]

    Mental labour

    Wouter Kool and Matthew Botvinick. Mental labour. Nature human behaviour, 2(12):899–908, 2018

  49. [57]

    Teaching categories to human learners with visual explanations

    Oisin Mac Aodha, Shihan Su, Yuxin Chen, Pietro Perona, and Yisong Yue. Teaching categories to human learners with visual explanations. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3820–3828, 2018

  50. [58]

    Xiao Dong and Caroline C. Hayes. Uncertainty visualizations: Helping decision makers become more aware of uncertainty and its implications. Journal of Cognitive Engineering and Decision Making, 6(1):30–56, 2012

  51. [59]

    Glassman, Jeremy Scott, Rishabh Singh, Philip J

    Elena L. Glassman, Jeremy Scott, Rishabh Singh, Philip J. Guo, and Robert C. Miller. Overcode: Visualizing variation in student solutions to programming problems at scale. ACM Trans. Comput.-Hum. Interact., 22(2), March 2015

  52. [60]

    Fumeng Yang, Zhuanyi Huang, Jean Scholtz, and Dustin L. Arendt. How do visual explanations foster end users’ appropriate trust in machine learning? In Proceedings of the 25th International Conference on Intelligent User Interfaces, IUI ’20, page 189–201, New York, NY , USA, 20...

  53. [61]

    Exploring the effect of explanation content and format on user comprehension and trust

    Antonio Rago, Bence Palfi, Purin Sukpanichnant, Hannibal Nabli, Kavyesh Vivek, Olga Kostopoulou, James Kinross, and Francesca Toni. Exploring the effect of explanation content and format on user comprehension and trust. arXiv preprint arXiv:2408.17401, 2024

  54. [62]

    Facilitating human-llm collaboration through factuality scores and source attribu- tions

    Hyo Jin Do, Rachel Ostrand, Justin D Weisz, Casey Dugan, Prasanna Sattigeri, Dennis Wei, Keerthiram Mu- rugesan, and Werner Geyer. Facilitating human-llm collaboration through factuality scores and source attribu- tions. arXiv preprint arXiv:2405.20434, 2024

  55. [63]

    A diachronic perspective on user trust in ai under uncertainty

    Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, and Mrinmaya Sachan. A diachronic perspective on user trust in ai under uncertainty. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5567–5580, 2023

  56. [64]

    Human feedback is not gold standard

    Tom Hosking, Phil Blunsom, and Max Bartolo. Human feedback is not gold standard. In The Twelfth Interna- tional Conference on Learning Representations, 2024

  57. [65]

    Display signaling in augmented reality: Effects of cue reliability and image realism on attention allocation and trust calibration

    Michelle Yeh and Christopher D Wickens. Display signaling in augmented reality: Effects of cue reliability and image realism on attention allocation and trust calibration. Human Factors, 43(3):355–365, 2001

  58. [66]

    When confidence meets accuracy: Exploring the effects of multiple perfor- mance indicators on trust in machine learning models

    Amy Rechkemmer and Ming Yin. When confidence meets accuracy: Exploring the effects of multiple perfor- mance indicators on trust in machine learning models. In Proceedings of the 2022 chi conference on human factors in computing systems, pages 1–14, 2022

  59. [67]

    The role of decision confidence in advice-taking and trust formation

    Niccolò Pescetelli and Nicholas Yeung. The role of decision confidence in advice-taking and trust formation. Journal of Experimental Psychology: General, 150(3):507, 2021. 33 Engaging with AI

  60. [68]

    are you really sure?

    Shuai Ma, Xinru Wang, Ying Lei, Chuhan Shi, Ming Yin, and Xiaojuan Ma. “are you really sure?” understand- ing the effects of human self-confidence calibration in ai-assisted decision making. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–20, 2024

  61. [69]

    What is ai literacy? competencies and design considerations

    Duri Long and Brian Magerko. What is ai literacy? competencies and design considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI ’20, page 1–16, New York, NY , USA,

  62. [70]

    Untan- gling critical interaction with ai in students’ written assessment

    Antonette Shibani, Simon Knight, Kirsty Kitto, Ajanie Karunanayake, and Simon Buckingham Shum. Untan- gling critical interaction with ai in students’ written assessment. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–6, 2024

  63. [71]

    A slow algorithm improves users’ assessments of the algorithm’s accuracy

    Joon Sung Park, Rick Barber, Alex Kirlik, and Karrie Karahalios. A slow algorithm improves users’ assessments of the algorithm’s accuracy. Proc. ACM Hum.-Comput. Interact., 3(CSCW), November 2019

  64. [72]

    Varshney, Amit Dhurandhar, and Richard Tomsett

    Charvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R. Varshney, Amit Dhurandhar, and Richard Tomsett. De- ciding fast and slow: The role of cognitive biases in ai-assisted decision-making. Proc. ACM Hum.-Comput. Interact., 6(CSCW1), April 2022

  65. [73]

    Harder, better, faster, stronger: Interactive visualiza- tion for human-centered ai tools

    Md Naimul Hoque, Sungbok Shin, and Niklas Elmqvist. Harder, better, faster, stronger: Interactive visualiza- tion for human-centered ai tools. arXiv preprint arXiv:2404.02147, 2024

  66. [74]

    From explainable to interactive ai: A literature review on current trends in human-ai interaction

    Muhammad Raees, Inge Meijerink, Ioanna Lykourentzou, Vassilis-Javed Khan, and Konstantinos Papange- lis. From explainable to interactive ai: A literature review on current trends in human-ai interaction. ArXiv, abs/2405.15051, 2024

  67. [75]

    Designing ai support for human involvement in ai-assisted decision making: A taxonomy of human-ai interactions from a systematic review

    Catalina Gomez, Sue Min Cho, Chien-Ming Huang, and Mathias Unberath. Designing ai support for human involvement in ai-assisted decision making: A taxonomy of human-ai interactions from a systematic review. arXiv preprint arXiv:2310.19778, 2023

  68. [76]

    Improving human-ai collaboration with descriptions of ai behavior

    Ángel Alexander Cabrera, Adam Perer, and Jason I Hong. Improving human-ai collaboration with descriptions of ai behavior. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1):1–21, 2023

  69. [77]

    The human-ai relationship in decision-making: Ai explanation to support people on justifying their decisions

    Juliana Jansen Ferreira and Mateus Monteiro. The human-ai relationship in decision-making: Ai explanation to support people on justifying their decisions. arXiv preprint arXiv:2102.05460, 2021

  70. [78]

    Ramos, Christopher Meek, Patrice Y

    Gonzalo A. Ramos, Christopher Meek, Patrice Y . Simard, Jina Suh, and Soroush Ghorashi. Interactive machine teaching: a human-centered approach to building machine-learned models. Human–Computer Interaction, 35:413 – 451, 2020

  71. [79]

    Fred Hohman, Andrew Head, Rich Caruana, Robert DeLine, and Steven M. Drucker. Gamut: A design probe to understand how data scientists understand machine learning models. In Proceedings of the 2019 CHI Confer- ence on Human Factors in Computing Systems, CHI ’19, page 1–13, New ...

  72. [80]

    Survey on visual analysis of event sequence data

    Yi Guo, Shunan Guo, Zhuochen Jin, Smiti Kaul, David Gotz, and Nan Cao. Survey on visual analysis of event sequence data. IEEE Transactions on Visualization and Computer Graphics, 28:5091–5112, 2020

  73. [81]

    Russell, and Aaron Hertzmann

    Zoya Bylinskii, Nam Wook Kim, Peter O’Donovan, Sami Alsheikh, Spandan Madan, Hanspeter Pfister, Frédo Durand, Bryan C. Russell, and Aaron Hertzmann. Learning visual importance for graphic designs and data visualizations. Proceedings of the 30th Annual ACM Symposium on User Int...

  74. [82]

    Are explanations helpful? a comparative study of the effects of explanations in ai- assisted decision-making

    Xinru Wang and Ming Yin. Are explanations helpful? a comparative study of the effects of explanations in ai- assisted decision-making. In Proceedings of the 26th International Conference on Intelligent User Interfaces , IUI ’21, page 318–328, New York, NY , USA, 2021. Associat...

  75. [83]

    Can users cor- rectly interpret machine learning explanations and simultaneously identify their limitations? arXiv preprint arXiv:2309.08438, 2023

    Yueqing Xuan, Edward Small, Kacper Sokol, Danula Hettiachchi, and Mark Sanderson. Can users cor- rectly interpret machine learning explanations and simultaneously identify their limitations? arXiv preprint arXiv:2309.08438, 2023

  76. [84]

    Are visual explanations useful? a case study in model-in-the-loop prediction

    Eric Chu, Deb Roy, and Jacob Andreas. Are visual explanations useful? a case study in model-in-the-loop prediction. arXiv preprint arXiv:2007.12248, 2020

  77. [85]

    Shiye Cao, Anqi Liu, and Chien-Ming Huang. Designing for appropriate reliance: Designing for appropriate reliance: The roles of ai uncertainty presentation, initial user decision, and user demographics in ai-assisted decision-making. arXiv preprint arXiv:2401.05612, 2024

  78. [86]

    Toward transparent ai: A survey on interpreting the inner structures of deep neural networks

    Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. Toward transparent ai: A survey on interpreting the inner structures of deep neural networks. In 2023 ieee conference on secure and trustworthy machine learning (satml), pages 464–483. IEEE, 2023. 34 Engaging with AI

  79. [87]

    Study on the helpfulness of explainable artificial intelligence

    Tobias Labarta, Elizaveta Kulicheva, Ronja Froelian, Christian Geißler, Xenia Melman, and Julian V on Klitzing. Study on the helpfulness of explainable artificial intelligence. In World Conference on Explainable Artificial Intelligence, pages 294–312. Springer, 2024

  80. [88]

    Katrin Glinka and Claudia Müller-Birn. Critical-reflective human-ai collaboration: Exploring computational tools for art historical image retrieval.Proceedings of the ACM on Human-Computer Interaction, 7(CSCW2):1– 33, 2023

  81. [89]

    The metacognitive demands and opportunities of generative ai

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. The metacognitive demands and opportunities of generative ai. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , CHI ’24, New Yo...

  82. [90]

    Manipulating and measuring model interpretability

    Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Wortman Vaughan, and Hanna Wallach. Manipulating and measuring model interpretability. InProceedings of the 2021 CHI conference on human factors in computing systems, pages 1–52, 2021

  83. [91]

    Understanding the effect of out-of-distribution examples and inter- active explanations on human-ai decision making

    Han Liu, Vivian Lai, and Chenhao Tan. Understanding the effect of out-of-distribution examples and inter- active explanations on human-ai decision making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1–45, 2021

  84. [92]

    A critical analysis of cognitive load measurement methods for evaluating the usability of different types of interfaces: guidelines and framework for human-computer interaction

    Ali Darejeh, Nadine Marcusa, Gelareh Mohammadi, and John Sweller. A critical analysis of cognitive load measurement methods for evaluating the usability of different types of interfaces: guidelines and framework for human-computer interaction. arXiv preprint arXiv:2402.11820, 2024

  85. [93]

    Emerging reliance behaviors in human-ai text generation: Hallucinations, data quality assessment, and cognitive forcing functions

    Zahra Ashktorab, Qian Pan, Werner Geyer, Michael Desmond, Marina Danilevsky, James M Johnson, Casey Dugan, and Michelle Bachman. Emerging reliance behaviors in human-ai text generation: Hallucinations, data quality assessment, and cognitive forcing functions. arXiv preprint ar...

  86. [94]

    Adaptive cognitive mechanisms to maintain calibrated trust and reliance in automation

    Christian Lebiere, Leslie M Blaha, Corey K Fallon, and Brett Jefferson. Adaptive cognitive mechanisms to maintain calibrated trust and reliance in automation. Frontiers in Robotics and AI, 8:652776, 2021

  87. [95]

    Evaluating the utility of conformal prediction sets for ai-advised image labeling

    Dongping Zhang, Angelos Chatzimparmpas, Negar Kamali, and Jessica Hullman. Evaluating the utility of conformal prediction sets for ai-advised image labeling. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–19, 2024

  88. [96]

    Using ai uncertainty quantification to improve human decision-making

    Laura Marusich, Jonathan Bakdash, Yan Zhou, and Murat Kantarcioglu. Using ai uncertainty quantification to improve human decision-making. In Forty-first International Conference on Machine Learning, 2023

  89. [97]

    The effects of over-reliance on ai dialogue systems on students’ cognitive abilities: a systematic review

    Chunpeng Zhai, Santoso Wibowo, and Lily D Li. The effects of over-reliance on ai dialogue systems on students’ cognitive abilities: a systematic review. Smart Learning Environments, 11(1):28, 2024

  90. [98]

    Automation bias: a systematic review of frequency, effect mediators, and mitigators

    Kate Goddard, Abdul Roudsari, and Jeremy C Wyatt. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1):121–127, 2012

  91. [99]

    Lasecki, Daniel S

    Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S. Lasecki, Daniel S. Weld, and Eric Horvitz. Beyond accuracy: The role of mental models in human-ai team performance. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):2–11, Oct. 2019

  92. [100]

    Learning argumentation skills through the use of prompts for self- explaining examples

    Silke Schworm and Alexander Renkl. Learning argumentation skills through the use of prompts for self- explaining examples. Journal of Educational Psychology, 99(2):285, 2007

  93. [101]

    The metacognitive demands and opportunities of generative ai

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. The metacognitive demands and opportunities of generative ai. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–24, 2024

  94. [102]

    Exploring the de- sign space of cognitive engagement techniques with ai-generated code for enhanced learning

    Majeed Kazemitabaar, Oliver Huang, Sangho Suh, Austin Z Henley, and Tovi Grossman. Exploring the de- sign space of cognitive engagement techniques with ai-generated code for enhanced learning. arXiv preprint arXiv:2410.08922, 2024

  95. [103]

    Generating situated reflection triggers about alternative solution paths: A case study of generative ai for computer-supported collaborative learning

    Atharva Naik, Jessica Ruhan Yin, Anusha Kamath, Qianou Ma, Sherry Tongshuang Wu, Charles Murray, Christopher Bogart, Majd Sakr, and Carolyn P Rose. Generating situated reflection triggers about alternative solution paths: A case study of generative ai for computer-supported co...

  96. [104]

    Comparing zealous and restrained ai recommendations in a real-world human-ai collaboration task

    Chengyuan Xu, Kuo-Chin Lien, and Tobias Höllerer. Comparing zealous and restrained ai recommendations in a real-world human-ai collaboration task. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–15, 2023

  97. [105]

    Statistical power analysis for the behavioral sciences

    Jacob Cohen. Statistical power analysis for the behavioral sciences. routledge, 2013. 35 Engaging with AI

  98. [106]

    Note on the sampling error of the difference between correlated proportions or percentages

    Quinn McNemar. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2):153–157, 1947

  99. [107]

    S. S. SHAPIRO and M. B. WILK. An analysis of variance test for normality (complete samples)†. Biometrika, 52(3-4):591–611, 12 1965

  100. [108]

    Individual comparisons by ranking methods

    Frank Wilcoxon. Individual comparisons by ranking methods. In Breakthroughs in statistics: Methodology and distribution, pages 196–202. Springer, 1992

  101. [109]

    Use of ranks in one-criterion variance analysis.Journal of the American statistical Association, 47(260):583–621, 1952

    William H Kruskal and W Allen Wallis. Use of ranks in one-criterion variance analysis.Journal of the American statistical Association, 47(260):583–621, 1952

  102. [110]

    A mathematical theory of communication

    Claude Elwood Shannon. A mathematical theory of communication. The Bell system technical journal , 27(3):379–423, 1948

  103. [111]

    Bignav: Bayesian infor- mation gain for guiding multiscale navigation

    Wanyu Liu, Rafael Lucas D’Oliveira, Michel Beaudouin-Lafon, and Olivier Rioul. Bignav: Bayesian infor- mation gain for guiding multiscale navigation. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems , CHI ’17, page 5869–5880, New York, NY , USA, ...

  104. [112]

    Wobbrock

    Mingrui Ray Zhang, Shumin Zhai, and Jacob O. Wobbrock. Text entry throughput: Towards unifying speed and accuracy in a single performance metric. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI ’19, page 1–13, New York, NY , USA, 2019. Asso...

  105. [113]

    Physical activity/exercise and diabetes: a position statement of the american diabetes association

    Sheri R Colberg, Ronald J Sigal, Jane E Yardley, Michael C Riddell, David W Dunstan, Paddy C Dempsey, Edward S Horton, Kristin Castorino, and Deborah F Tate. Physical activity/exercise and diabetes: a position statement of the american diabetes association. Diabetes care, 39(1...

  106. [114]

    Nutrition therapy for adults with diabetes or prediabetes: a consensus report

    Alison B Evert, Michelle Dennison, Christopher D Gardner, W Timothy Garvey, Ka Hei Karen Lau, Janice MacLeod, Joanna Mitri, Raquel F Pereira, Kelly Rawlings, Shamera Robinson, et al. Nutrition therapy for adults with diabetes or prediabetes: a consensus report. Diabetes care, ...

  107. [2020]

    Association for Computing Machinery

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.