REVIEW 3 major objections 49 references
Users’ preferred explanation styles for AI privacy redactions change with domain and how much is redacted, and trust is higher when they get the styles they choose.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 23:23 UTC pith:BT2UTOG2
load-bearing objection Solid free-choice HCI study on explanation styles under privacy redaction; the preference–trust claim is confounded by choice itself, but the context and individual-difference patterns still stand. the 3 major comments →
Exploring the Interaction of Explanation Styles, Context, and Trust of AI Privacy Redaction in AI-mediated Interactions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Explanation preferences for AI privacy redaction are not fixed: they vary systematically with domain and redaction level (and their interaction), users’ trust in the mediator is higher when they receive the explanation style(s) they prefer than when they receive no explanation or a randomly assigned one, and individual users differ both in which styles they favor and in how strongly context drives their choices.
What carries the argument
A four-stage LLM pipeline (information generation for a domain and redaction level, information redaction, generation of five named explanation styles—contrastive, general, thorough, normative, causal—and subsequent redaction of the explanations themselves) paired with a mixed-design user study that lets participants choose preferred styles and rate trust.
Load-bearing premise
The five LLM-generated explanation styles are different enough from one another that people’s choices reflect real style preferences rather than surface wording or residual confusion between styles.
What would settle it
A follow-up study that re-uses the same scenarios but forces every participant to receive only one randomly assigned style (or none) and finds no trust difference relative to free choice, or that finds no reliable domain or redaction-level differences in free-choice frequencies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how explanation style preferences and trust interact with context (domain and redaction amount) in AI-mediated privacy redaction. An LLM pipeline generates domain-specific content at three redaction levels, redacts sensitive items, produces five explanation styles (contrastive, general, thorough, normative, causal), and redacts the explanations. A Prolific study (n=249) lets each participant choose preferred style(s) for six of eighteen scenarios and then rate trust. Results claim that preferences vary systematically with domain and redaction level (and their interaction for general/normative styles), that trust is higher when users receive chosen explanations than in a prior forced/no-explanation study, that individuals differ in context dependence of preferences, that age/education and baseline AI trust correlate modestly with choices/trust, and that users prefer low-effort feedback (categorizing/sorting).
Significance. If the preference–context and preference–trust findings hold under cleaner controls, the work supplies concrete design guidance for adaptive, personalized explanation interfaces in privacy-sensitive AI mediation—an increasingly relevant HCI setting. Strengths include a multi-domain design, power analysis, verification of redaction completeness (manual + LLM judge), an LLM-judge style-distinctness check, and a usable context-dependence score. The contribution is primarily empirical and design-oriented rather than theoretical; its value depends on whether the trust increment can be attributed to style match rather than choice agency, and on whether the five styles are sufficiently distinct and ecologically valid.
major comments (3)
- §5.2.2 / Fig. 7 / RQ2: The central claim that “users’ trust is linked to preferred explanation” rests on comparing Group 3 (current free-choice study) to Groups 1–2 from the authors’ prior forced-assignment study (no explanation or random general/thorough). Participants in Group 3 always choose preferred style(s) before rating trust, so the design confounds style–preference match with the act of choosing (agency/control) and with demand characteristics of the free-choice interface. There is no within-study arm that forces a non-preferred style after preference elicitation, nor a yoked control that assigns the same style without choice. The reported F(2,2223)=16.19, p<0.001 therefore cannot cleanly attribute the trust increment to style content. This comparison is load-bearing for the abstract and conclusion claims and needs either a controlled re-analysis/arm or a substantially narrowed
- §3.3 / Fig. 3: Style distinctness is only partially supported. The LLM-judge confusion matrix shows 0.65 overlap between causal and thorough (and non-trivial off-diagonals elsewhere). Preference differences and the co-occurrence matrix (§5.3.2) may therefore partly reflect surface wording rather than the named styles. Because the paper’s design implications treat the five styles as meaningfully different levers for adaptive interfaces, residual confusion weakens the claim that observed effects are style-specific. Human validation of style discriminability (or a reduced style set) is needed before the preference results can be interpreted as style effects.
- Abstract vs. body inconsistency: The abstract reports n=180, privacy-effectiveness effects (p<0.05, d≈0.3), and greater reliance on explanations under extensive redaction (f≈0.2). The body reports n=249, five RQs focused on free choice among five styles, and no significant trust differences by domain/redaction when preferred styles are given (§5.2.1). These are not the same study claims. The abstract must be rewritten to match the actual design, sample, measures, and results; otherwise readers cannot evaluate the contribution.
Circularity Check
Empirical HCI study; only minor self-citation load in the cross-study trust comparison, not definitional circularity.
specific steps
-
self citation load bearing
[§5.2.2 Explanation Choice vs. Trust; Fig. 7; citation [20]]
"To explore the takeaway further, we compare the trust values from a previous study [20] that explored trust using the same Likert-style measures as the current work. In that study, participants were randomly assigned into three conditions: no explanation, general explanations, and thorough explanations. We then grouped the data into three groups: Group 1: Prior study participants who did not receive explanations; Group 2: Prior study participants who received an explanation randomly (general or thorough); Group 3: Current study participants who were able to chose their desired explanation(s) …"
The central claim that preferred/chosen explanations raise trust rests on contrasting the current free-choice cohort against no-explanation and random-explanation arms taken exclusively from the authors’ own prior arXiv. The baseline is therefore not an external benchmark but a self-cited related experiment; the trust increment is load-bearing for the abstract and RQ2 takeaway yet is not independently re-measured under forced non-preferred assignment in the present design. This is mild self-citation load-bearing, not definitional circularity.
full rationale
This is a user-study paper (n=249) measuring explanation-style preferences and trust under domain × redaction-level conditions. The main results (RQ1 preference variation by context; RQ3 individual context-dependence scores; RQ4 demographics; RQ5 feedback-type rankings) are direct empirical outcomes of free choice and Likert ratings; they do not reduce by construction to the generation pipeline or to any fitted parameter. The five explanation styles are stipulated inputs (prompted from prior XAI literature), verified for distinctness by an LLM judge (Fig. 3), and then offered as choice options—preference frequencies are measured, not derived. The sole circularity-adjacent step is the RQ2 trust comparison (Fig. 7): Groups 1–2 (no / random explanation) are taken from the authors’ own prior arXiv [20] and contrasted with Group 3 (current free-choice condition). That is self-citation of a related experiment used as a baseline, not a uniqueness theorem or a definitional identity. The comparison is load-bearing for the claim that “chosen explanation raises trust,” but the circularity is mild (score 2): the prior study is an independent data collection, not a mathematical premise that forces the present result. No self-definitional loop, no fitted-input-called-prediction, and no ansatz smuggled via citation appear in the derivation chain. Residual style confusion (causal–thorough 0.65) and the free-choice confound are validity/correctness concerns, not circularity.
Axiom & Free-Parameter Ledger
free parameters (3)
- redaction-level thresholds (heavy 5–7, moderate 3–5, light 1–3 sensitive types)
- number of domains (6) and scenarios per participant (6 of 18)
- five explanation styles selected from literature
axioms (4)
- domain assumption LLM (GPT-5) redaction and explanation generation plus Gemma3 judging produce valid, privacy-safe stimuli whose style labels match the intended constructs
- domain assumption Self-report 5-pt Likert items adapted from Hoffman et al. and Jian et al. validly measure trust and explanation satisfaction in this setting
- domain assumption Prolific sample and six randomly assigned scenarios per person generalize to real AI-mediated privacy interactions
- standard math ANOVA/Tukey on choice proportions and mean trust are appropriate for the mixed design
invented entities (1)
-
context dependence score (normalized average CV of style proportions across domains/redaction levels)
no independent evidence
read the original abstract
AI-mediated communication is increasingly being utilized to help facilitate interactions; however, in privacy sensitive domains, an AI mediator has the additional challenge of considering how to preserve privacy. In these contexts, a mediator may redact or withhold information, raising questions about how users perceive these interventions and whether explanations of system behavior can improve trust. In this work, we investigate how explanations of redaction operations can affect user trust in AI-mediated communication. We devise a scenario where a validated system removes sensitive content from messages and generates explanations of varying detail to communicate its decisions to recipients. We then conduct a user study with 180 participants that studies how user trust and preferences vary for cases with different amounts of redacted content and different levels of explanation detail. Our results show that participants believed our system was more effective at preserving privacy when explanations were provided (p<0.05, Cohen's d ~ 0.3). We also found that contextual factors had an impact; participants relied more on explanations and found them more helpful when the system performed extensive redactions (p<0.05, Cohen's f ~ 0.2). We also found that explanation preferences depended on individual differences as well, and factors such as age and baseline familiarity with AI affected user trust in our system. These findings highlight the importance and challenge of balancing transparency and privacy in AI-mediated communications and suggest that adaptive, context-aware explanations are essential for designing privacy-aware, trustworthy AI systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Emmanuel A. Abbe, Amir E. Khandani, and Andrew W. Lo. 2012. Privacy- preserving methods for sharing financial risk exposures.American Economic Review102, 3 (2012), 65–70. ISBN: 0002-8282 Publisher: American Economic Association
work page 2012
-
[2]
Ashraf Abdul, Jo Vermeulen, Danding Wang, Brian Y. Lim, and Mohan Kankan- halli. 2018. Trends and trajectories for explainable, accountable and intelligible Exploring the Interaction of Explanation Styles, Context, and Trust of AI Privacy Redaction in AI-mediated Interactions UIST ’26, November 2026, Detroit, MI systems: An hci research agenda. InProceedi...
work page 2018
-
[3]
Amina Adadi and Mohammed Berrada. 2018. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI).IEEE access6 (2018), 52138– 52160. ISBN: 2169-3536
work page 2018
-
[4]
Jorre, Claudia Pagliari, Ruth Jepson, and Sarah Cunningham-Burley
Mhairi Aitken, Jenna de St. Jorre, Claudia Pagliari, Ruth Jepson, and Sarah Cunningham-Burley. 2016. Public responses to the sharing and linkage of health data for research purposes: a systematic review and thematic synthesis of quali- tative studies.BMC medical ethics17, 1 (2016), 73. ISBN: 1472-6939 Publisher: Springer
work page 2016
-
[5]
Saleema Amershi, Maya Cakmak, William Bradley Knox, and Todd Kulesza. 2014. Power to the people: The role of humans in interactive machine learning.AI magazine35, 4 (2014), 105–120. ISBN: 2371-9621
work page 2014
-
[6]
Sílvia Araújo and Micaela Aguiar. 2023. Simplifying specialized texts with AI: a ChatGPT-based learning scenario. InInternational Conference in Information Technology and Education. Springer, 599–609
work page 2023
-
[7]
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Ben- netot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2019. Explain- able Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Chal- lenges toward Responsible AI. doi:10.4...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1910.10045 2019
-
[8]
Cai, Jonas Jongejan, and Jess Holbrook
Carrie J. Cai, Jonas Jongejan, and Jess Holbrook. 2019. The effects of example- based explanations in a machine learning interface. InProceedings of the 24th international conference on intelligent user interfaces. 258–262
work page 2019
-
[9]
Kelly Caine and Rima Hanania. 2013. Patients want granular privacy control over health information in electronic medical records.Journal of the American Medical Informatics Association20, 1 (2013), 7–15. ISBN: 1067-5027
work page 2013
- [10]
-
[11]
Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of inter- pretable machine learning.arXiv preprint arXiv:1702.08608(2017)
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[12]
Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, and Mark O Riedl
Upol Ehsan, Samir Passi, Q. Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, and Mark O Riedl. 2024. The Who in XAI: How AI Background Shapes Perceptions of AI Explanations. InProceedings of the CHI Conference on Human Factors in Computing Systems. ACM, Honolulu HI USA, 1–32. doi:10.1145/3613904.3642474
-
[13]
Ghada El Haddad, Esma Aimeur, and Hicham Hage. 2018. Understanding trust, privacy and financial fears in online payment. In2018 17th IEEE international con- ference on trust, security and privacy in computing and communications/12th IEEE international conference on big data science and engineering (TrustCom/BigDataSE). IEEE, 28–36
work page 2018
-
[14]
Tesca Fitzgerald, Pallavi Koppol, Patrick Callaghan, Russell Quinlan Jun Hei Wong, Reid Simmons, Oliver Kroemer, and Henny Admoni. 2022. Inquire: Inter- active querying for user-aware informative reasoning. In6th Annual Conference on Robot Learning
work page 2022
-
[15]
A. K. M. Haque. 2025. Explainable artificial intelligence (XAI): making AI under- standable for end users. (2025). ISBN: 9524122391
work page 2025
-
[16]
Metrics for Explainable AI: Challenges and Prospects
Robert R. Hoffman, Shane T. Mueller, Gary Klein, and Jordan Litman. 2019. Metrics for Explainable AI: Challenges and Prospects. doi:10.48550/arXiv.1812.04608 arXiv:1812.04608 [cs]
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1812.04608 2019
-
[17]
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. 2021. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 624–635
work page 2021
-
[18]
Jiun-Yin Jian, Ann M. Bisantz, and Colin G. Drury. 2000. Foundations for an empirically determined scale of trust in automated systems.International journal of cognitive ergonomics4, 1 (2000), 53–71. ISBN: 1088-6362
work page 2000
-
[19]
Linda A. Jones, Jenny R. Nelder, Joseph M. Fryer, Philip H. Alsop, Michael R. Geary, Mark Prince, and Rudolf N. Cardinal. 2022. Public opinion on sharing data from health services for clinical and research purposes without explicit consent: an anonymous online survey in the UK.Bmj Open12, 4 (2022), e057579. ISBN: 2044-6055 Publisher: British Medical Journ...
work page 2022
- [20]
-
[21]
Katherine C. Kellogg, Melissa A. Valentine, and Angele Christin. 2020. Algorithms at work: The new contested terrain of control.Academy of management annals 14, 1 (2020), 366–410. ISBN: 1941-6520 Publisher: Briarcliff Manor, NY
work page 2020
-
[22]
Sunnie SY Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. " help me help the ai": Understanding how explainability can support human-ai interaction. Inproceedings of the 2023 CHI conference on human factors in computing systems. 1–17
work page 2023
-
[23]
René F. Kizilcec. 2016. How Much Information?: Effects of Transparency on Trust in an Algorithmic Interface. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems. ACM, San Jose California USA, 2390–2395. doi:10.1145/2858036.2858402
-
[24]
Rafal Kocielnik, Saleema Amershi, and Paul N. Bennett. 2019. Will You Accept an Imperfect AI?: Exploring Designs for Adjusting End-user Expectations of AI Systems. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM, Glasgow Scotland Uk, 1–14. doi:10.1145/3290605.3300641
-
[25]
Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of explanatory debugging to personalize interactive machine learning. InProceedings of the 20th international conference on intelligent user interfaces. 126–137
work page 2015
-
[26]
Todd Kulesza, Simone Stumpf, Margaret Burnett, Sherry Yang, Irwin Kwan, and Weng-Keen Wong. 2013. Too much, too little, or just right? Ways explanations impact end users’ mental models. In2013 IEEE Symposium on visual languages and human centric computing. IEEE, 3–10
work page 2013
-
[27]
Gershman, and Finale Doshi-Velez
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. 2019. Human evaluation of models built for interpretability. InProceedings of the AAAI conference on human computation and crowdsourcing, Vol. 7. 59–67
work page 2019
-
[28]
Retno Larasati, Anna De Liddo, and Enrico Motta. 2020. The effect of explana- tion styles on user’s trust. InProceedings of IUI workshop on Explainable Smart Systems and Algorithmic Transparency in Emerging Technologies, Vol. 2582. CEUR Workshop Proceedings
work page 2020
-
[29]
John Lee and Neville Moray. 1992. Trust, control strategies and allocation of function in human-machine systems.Ergonomics35, 10 (1992), 1243–1270. ISBN: 0014-0139
work page 1992
-
[30]
John D. Lee and Katrina A. See. 2004. Trust in automation: Designing for appro- priate reliance.Human factors46, 1 (2004), 50–80. ISBN: 0018-7208
work page 2004
-
[31]
Tianshi Li, Sauvik Das, Hao-Ping Lee, Dakuo Wang, Bingsheng Yao, and Zhiping Zhang. 2024. Human-Centered Privacy Research in the Age of Large Language Models. doi:10.48550/arXiv.2402.01994 arXiv:2402.01994 [cs]
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2402.01994 2024
-
[32]
Brian Y. Lim, Anind K. Dey, and Daniel Avrahami. 2009. Why and why not explanations improve the intelligibility of context-aware intelligent systems. InProceedings of the SIGCHI conference on human factors in computing systems. 2119–2128
work page 2009
-
[33]
Tim Miller. 2019. Explanation in artificial intelligence: Insights from the social sciences.Artificial intelligence267 (2019), 1–38. ISBN: 0004-3702
work page 2019
-
[34]
Helen Nissenbaum. 2004. Privacy as Contextual Integrity.Washington Law Review79 (2004)
work page 2004
- [35]
-
[36]
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wort- man Wortman Vaughan, and Hanna Wallach. 2021. Manipulating and measuring model interpretability. InProceedings of the 2021 CHI conference on human factors in computing systems. 1–52
work page 2021
-
[37]
Priscilla M. Regan and Jolene Jesse. 2019. Ethical challenges of edtech, big data and personalized learning: Twenty-first century student sorting and tracking.Ethics and Information Technology21, 3 (2019), 167–179. ISBN: 1388-1957 Publisher: Springer
work page 2019
-
[38]
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1135–1144
work page 2016
-
[39]
K. Saravanakumar and K. Deepa. 2016. On privacy and security in social media–a comprehensive study.Procedia computer science78 (2016), 114–119. ISBN: 1877-0509 Publisher: Elsevier
work page 2016
- [40]
-
[41]
Aditya Singhal, Nikita Neveditsin, Hasnaat Tanveer, and Vijay Mago. 2024. To- ward fairness, accountability, transparency, and ethics in AI for social media and health care: scoping review.JMIR medical informatics12, 1 (2024), e50048. Publisher: JMIR Publications Inc., Toronto, Canada
work page 2024
-
[42]
Sharon Slade and Paul Prinsloo. 2013. Learning analytics: Ethical issues and dilemmas.American behavioral scientist57, 10 (2013), 1510–1529. ISBN: 0002-7642 Publisher: SAGE Publications Sage CA: Los Angeles, CA
work page 2013
-
[43]
Mena Teebken and Thomas Hess. 2021. Privacy in a digitized workplace: Towards an understanding of employee privacy concerns. (2021)
work page 2021
-
[44]
Sahil Verma, Varich Boonsanong, Minh Hoang, Keegan Hines, John Dickerson, and Chirag Shah. 2024. Counterfactual explanations and algorithmic recourses for machine learning: A review.Comput. Surveys56, 12 (2024), 1–42. ISBN: 0360-0300
work page 2024
-
[45]
As an AI language model, I cannot
Joel Wester, Tim Schrills, Henning Pohl, and Niels van Berkel. 2024. “As an AI language model, I cannot”: Investigating LLM Denials of User Requests. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–14. doi:10.1145/3613904.3642135
-
[46]
Afrizal Zein. 2025. Personalized Explainable AI: Dynamic Adjustment of Ex- planations for Novice and Expert Users.Jurnal Penelitian Pendidikan IPA11, 9 (2025), 513–520. ISBN: 2407-795X. UIST ’26, November 2026, Detroit, MI Kaushik, et. al
work page 2025
-
[47]
Shuning Zhang, Ying Ma, Jingruo Chen, Simin Li, Xin Yi, and Hewu Li. 2025. Towards Aligning Personalized Conversational Recommendation Agents with Users’ Privacy Preferences. doi:10.48550/arXiv.2508.07672 arXiv:2508.07672 [cs]
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2508.07672 2025
-
[48]
Mingqian Zheng, Wenjia Hu, Patrick Zhao, Motahhare Eslami, Jena D. Hwang, Faeze Brahman, Carolyn Rose, and Maarten Sap. 2025. Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences. doi:10. 48550/arXiv.2506.00195 arXiv:2506.00195 [cs]. A Appendix Exploring the Interaction of Explanation Styles, Context, and Trust of A...
-
[49]
Parsed text; matched handles (@...), emails, IPs, postal codes, device models, and city/state names. 2) Tagged as usernames, emails, IPs, locations, device IDs. 3) Generalized: accounts to participants; cities to metropolitan areas; zip/IP to non-specific ranges; emails to undisclosed; devices to smartphones. 4) Validated with a second pass: no identifiab...
work page 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.