REVIEW 5 major objections 5 minor 105 references
Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness
T0 review · 5 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Telling users an AI is superhuman reshapes trust and fairness judgments, but the effect flips with expertise and context.
desk verdict The persona prompts encode the heuristics the paper then 'discovers', and the N=32 human data contradict the fairness claim; the paper is candid about this, but the abstract overclaims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Persona Cards, a standardized specification of a synthetic user inspired by Model Cards: each card fixes the persona's skill level, strategic preferences, named cognitive heuristics (e.g., availability heuristic, anchoring, confirmation bias, authority bias), belief-updating rules, and an example response, and is instantiated as an LLM prompt that generates Likert ratings and open-text justifications for toxicity, fairness, and trust. The cards carry the argument by making the simulated user population transparent, diverse, and reproducible, and by seeding the human validation study so the same disclosure conditions can be compared across synthetic and real users.
What would settle it
A preregistered replication with real novices and experts (e.g., 50 per condition) playing the same capped-APM StarCraft II AI under the three disclosure conditions could settle the claim: if novices do not show the predicted rise in trust under superhuman disclosure, or if experts do not show the predicted drop in toxicity when the AI is disclosed as superhuman, the central interaction claim would be contradicted; the paper's own human fairness data already run against the persona prediction.
Extended reading notes
Core claim
The paper's central claim is that capability disclosure is an interpretive frame: it changes what a user's experience of an AI means, and the direction of that change depends on the user's expertise and on whether the AI is an opponent or an assistant. In StarCraft II, personas and a 32-person human validation both indicated that an AI revealed as superhuman was seen as less toxic than a deceptively human-level one, but the same disclosure raised trust among novices and intermediates to the point of overreliance, while experts began treating the AI as unbeatable and switched to sub-goals such as prolonging the match. In cooperative chatbot interactions, novices and intermediates rated a disclosed 'superhuman' assistant more trustworthy and fairer, whereas experts found the disclosure repetitive and rated the assistant more toxic and less fair. The paper concludes that transparency is not a cure-all; its effects are systematically moderated by expertise and context, so disclosure must be tailored rather than applied uniformly.
Load-bearing premise
The load-bearing premise is that LLM-generated personas, prompted to enact specified cognitive heuristics, produce ratings and behaviors representative enough of real users to draw conclusions about human reactions; the paper's own 32-person validation shows only partial and sometimes opposite agreement (notably in fairness ratings), so that premise is not yet established.
Editorial extensions
If this is right
- Game developers can use capability disclosure to soften accusations of cheating, since both personas and human players rated a disclosed superhuman AI as less toxic than an undisclosed one.
- Disclosure without safeguards is risky for novices: the large trust increases under disclosure point to overreliance on a system presented as infallible.
- For experts, flat disclosure can backfire on performance, because it converts a beatable-looking opponent into an 'unbeatable' one and pushes players toward sub-goals like stalling the match.
- In cooperative AI assistants, the same disclosure text is reassuring to novices but patronizing to experts, so user-adaptive disclosure strategies are needed rather than a single transparency statement.
Reading between the lines
- Beyond the paper: the fairness ratings in the human validation ran opposite to the persona predictions for novices and intermediates, so the persona results may overstate how generalizable the expertise-by-disclosure interaction is; a larger human sample could reverse the flip.
- Beyond the paper: the same double-edged dynamic should appear in non-game domains where an AI's capability claim can be checked against visible errors, such as medical or legal advice, where experts may under-rely after a 'superhuman' label fails to match observed mistakes.
- Beyond the paper: a testable extension is to vary the wording of disclosure (e.g., 'superhuman' vs 'highly skilled') to see whether the framing effect has an optimum, since the paper shows the label itself, not just the information, carries the effect.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how disclosing that a game AI is superhuman affects users' trust, fairness perceptions, and toxicity perceptions, and whether user expertise moderates these effects. The primary evidence comes from LLM-driven synthetic personas ('Persona Cards') assigned explicit cognitive heuristics, supplemented by a small human validation study (N=32) in StarCraft II and a separate LLM-chat scenario reported in an appendix. The authors report that disclosure reduces toxicity and increases trust in competitive play, but that its effects on fairness and reliance are context- and expertise-dependent; they release the Persona Cards Dataset and propose design guidelines for adaptive transparency.
Significance. The research question—whether disclosing superhuman AI capability helps or harms user trust and perceived fairness—is timely and practically important, and the paper's effort to release prompts, logs, and protocols is commendable for reproducibility. If the effects were empirically established, the findings would usefully inform transparency guidelines for game AI and conversational agents. However, the central empirical claim is not established: the primary results are generated by personas whose prompts explicitly encode the very heuristics the paper reports as findings, the data are filtered to discard runs that deviate from those heuristics, and the only independent human data (N=32) directly contradict the headline fairness result. Because the load-bearing bridge from synthetic personas to claims about human users is unverified, the paper's contribution is currently a methodological proposal plus a dataset, not a validated empirical demonstration of transparency effects.
major comments (5)
- [§3.2, Appendix B] The personas are constructed by assigning specific cognitive heuristics (e.g., Authority Bias, Confirmation Bias) in the persona cards (Tables 2, 3, 6–14), and the prompts explicitly instruct the model to apply them. For example, Appendix B.2's LLM_Novice_1 prompt says, 'Considering your persona's tendency to accept LLM responses as factual... how does this information affect your perception?' The reported findings that novices overtrust and experts are skeptical are therefore close to a restatement of the prompt design. The abstract's claim that the results 'reveal' these effects overstates what can be learned from the synthetic-persona pipeline.
- [§7 (Ethical Considerations and Limitations)] The consistency-checking procedure discards and regenerates any run in which a persona 'contradicted its stated traits' and repeats 'until the persona consistently followed its persona card attributes.' This is a selection bias that removes the very variation needed to test whether the assigned heuristics predict responses. Consequently, the p-values reported in §4.1 are computed on a filtered sample and are not interpretable as evidence about the population of persona responses, let alone about human users.
- [§5.1 (Grounded Data Analysis)] The human validation data (N=32) directly contradict the headline fairness result of the synthetic study. The paper reports that human novices rated Superhuman–No Disclosure fairness higher than Superhuman–Disclosure (M=3.50 vs. 2.20) and that intermediate players also rated No Disclosure higher (M=4.25 vs. 2.67), whereas §4.1 claims disclosure increases fairness for novices. No inferential statistics are reported for the human data, so the 'validated across both adversarial and collaborative contexts' claim in the contributions list is unsupported by the evidence presented.
- [§4.1.1 (Fairness paragraph)] The text is internally inconsistent. It first states 'Superhuman – Disclosure generated a decrease in Fairness (about -0.7, p < 0.05)' and then immediately states 'the actual ratings show that undisclosed superhuman skill was perceived as even less fair,' which contradicts the reported direction of the effect. Table 4 lists Fairness values '3.65 −0.7 3.45' with an interaction p<0.001, but the text does not reconcile these numbers with the surrounding prose.
- [§4.1, §5.1] The statistical reporting for the StarCraft II experiment is incomplete in ways that prevent verification. No sample sizes or degrees of freedom are given for the ANOVAs, several means are reported without standard deviations (e.g., Expert fairness M=3.45 in §4.1 and the human-participant means in §5.1), and the Mann-Whitney U tests are reported with U statistics but no per-group n. Without these details, the claimed effects and their magnitudes cannot be independently checked.
minor comments (5)
- [§5.2.2] There is an empty citation at 'the implicit social contract of competitive play []' that should be filled or removed.
- [§5.1] The phrase 'didn't necessitate any forms of aggression, such as taunting or jaundice' appears to contain a word error ('jaundice' for 'jeering'); please revise.
- [§5.1] The human-participant results are reported only as means and standard deviations, with no tests or effect sizes; for a paper whose central claim is empirical validation, this is a substantial presentation gap that should at least be acknowledged explicitly as a limitation.
- [Appendix D] The paper references 'Table 25' for the LLM experiment trust ANOVA, but the appendix tables are numbered 17–24; the reference should be updated to the correct table number.
- [§1 and §3.1] The abstract and introduction describe 'frustration and strategic defeatism among novices in cooperative scenarios,' but the reported cooperative (LLM) results in Appendix D are only about trust, toxicity, and fairness; the link between disclosure and 'strategic defeatism' is asserted rather than demonstrated with the presented data.
Circularity Check
Persona-card prompts prescribe the reported heuristics and non-conforming runs are discarded, so the central empirical disclosure effects are circular; the N=32 human fairness data run opposite to the synthetic claim.
-
self definitional
[Appendix D.2, Table 18 (LLM_Novice_1 Persona Card)]
"Authority Bias: The persona tends to accept the LLM’s responses as factual and authoritative without questioning their accuracy or validity. Confirmation Bias: The persona is more likely to believe information provided by the LLM if it aligns with their existing views or beliefs."
The novice overtrust and overreliance reported under disclosure is the same bias written into LLM_Novice_1’s card. Since the persona is defined as accepting LLM responses as factual, observing that this persona trusts the LLM is a restatement of the definition, not an empirical consequence of the disclosure manipulation.
-
other
[Section 7 (Ethical Considerations and Limitations)]
"If the persona’s answers indicated confusion or contradicted its stated traits (e.g., a “Novice” persona suddenly referencing expert-level strategies), we discarded that run and regenerated. We repeated this process until the persona consistently followed its persona card attributes."
This is an explicit selection-for-consistency step: any run that would have produced evidence against the persona card’s assumed heuristics is removed before the statistical analyses. The surviving corpus therefore cannot test whether novices are biased or experts are skeptical; it is filtered to satisfy exactly those assumptions.
1 more flagged steps
-
self definitional
[Appendix B.2, Prompt (LLM_Novice_1)]
"Considering your persona’s tendency to accept LLM responses as factual and to believe information aligning with existing views, how does this information affect your perception of the chatbot?"
The rating question is preceded by an instruction that tells the model which mechanism to apply. The model’s subsequent trust and toxicity ratings are therefore a direct fulfillment of a demand characteristic: the prompt names the expected heuristic and asks for a rating in light of it. The reported 'novice overtrust' is the prompt instruction re-read as observed behavior.
full rationale
The paper’s central claims—that disclosure increases trust and overreliance in novices and provokes skepticism in experts—are generated by LLM personas whose Persona Cards explicitly assign those very heuristics, and whose prompts instruct the model to answer rating questions by considering those heuristics. Section 7’s consistency filter then discards runs that contradict the stated traits, so the analyzed synthetic corpus cannot disconfirm the built-in expertise/heuristic pattern. The only independent check, the N=32 human study, does not rescue the claim: Section 5.1 reports novice human participants gave lower fairness scores to Superhuman–Disclosure (M = 2.20) than to Superhuman–No Disclosure (M = 3.50), the opposite of the synthetic fairness pattern, and the paper itself concedes the personas 'significantly exaggerated' toxicity ratings. The Persona Cards framework is a legitimate methodological contribution and the paper is transparent about limitations, but as an 'empirical demonstration' of disclosure effects the main results are circular: the outcome variable is largely written into the input prompt and the selection rule. Score 8 reflects that the central claim reduces by construction to its own inputs, not merely incidental self-citation.
Assumptions & free parameters
free parameters (4)
- Persona heuristic specifications
- APM ranges for skill tiers =
Novice 35-50; Expert 200+; Superhuman adds ~300
- Disclosure statement wording =
'far higher Actions Per Minute', 'markedly superior reasoning'
- Number of LLM runs per condition =
80 (inferred from ANOVA df 1,76)
assumptions (4)
- domain assumption LLM-generated personas are valid proxies for human perceptions
- domain assumption Cognitive heuristics assigned to personas are behaviorally realistic
- ad hoc to paper Consistency filtering does not bias results
- standard math Two-way ANOVA is valid despite normality violations
invented entities (1)
-
Persona Cards (synthetic personas)
Cite this review
Pith. "Pith review of Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness." pith.science (2026). https://pith.science/paper/OZKL3NEB
@misc{pith2026250315514,
author = {Pith},
title = {Pith review of: Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness},
year = {2026},
howpublished = {\url{https://pith.science/paper/OZKL3NEB}},
note = {Machine review of arXiv:2503.15514}
}
read the original abstract
As artificial intelligence surpasses human performance in select tasks, disclosing superhuman capabilities poses distinct challenges for fairness, accountability, and trust. However, the impact of such disclosures on diverse user attitudes and behaviors remains unclear, particularly concerning potential negative reactions like discouragement or overreliance. This paper investigates these effects by utilizing Persona Cards: a validated, standardized set of synthetic personas designed to simulate diverse user reactions and fairness perspectives. We conducted an ethics board-approved study (N=32), utilizing these personas to investigate how capability disclosure influenced behaviors with a superhuman game AI in competitive StarCraft II scenarios. Our results reveal transparency is double-edged: while disclosure could alleviate suspicion, it also provoked frustration and strategic defeatism among novices in cooperative scenarios, as well as overreliance in competitive contexts. Experienced and competitive players interpreted disclosure as confirmation of an unbeatable opponent, shifting to suboptimal goals. We release the Persona Cards Dataset, including profiles, prompts, interaction logs, and protocols, to foster reproducible research into human alignment AI design. This work demonstrates that transparency is not a cure-all; successfully leveraging disclosure to enhance trust and accountability requires careful tailoring to user characteristics, domain norms, and specific fairness objectives.
Figures
Reference graph
Works this paper leans on
-
[1]
Martin Adam, Michael Wessel, and Alexander Benlian. 2021. AI-based chatbots in customer service and their effects on user compliance. Electronic Markets 31, 2 (2021), 427–445
2021
-
[2]
NIST AI. 2023. Artificial intelligence risk management framework (AI RMF 1.0). URL: https://nvlpubs. nist. gov/nistpubs/ai/nist. ai (2023), 100–1
2023
-
[3]
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565 (2016)
arXiv 2016
-
[4]
Lisa P Argyle, Ethan C Busby, Nancy Fulda, Joshua R Gubler, Christopher Rytting, and David Wingate. 2023. Out of one, many: Using language models to simulate human samples. Political Analysis 31, 3 (2023), 337–351
2023
-
[5]
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems . 1–16
2021
-
[6]
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Fran- cis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al. 2020. The hanabi challenge: A new frontier for ai research.Artificial Intelligence 280 (2020), 103216
2020
-
[7]
Benoit Bediou, Melissa A Rodgers, Elizabeth Tipton, Richard E Mayer, C Shawn Green, and Daphne Bavelier. 2023. Effects of action video game play on cognitive skills: A meta-analysis. (2023)
2023
-
[8]
Nicole A Beres, Julian Frommel, Elizabeth Reid, Regan L Mandryk, and Madison Klarkowski. 2021. Don’t you know that you’re toxic: Normalization of toxicity in online gaming. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–15
2021
Show all 105 references
-
[9]
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al. 2019. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680 (2019)
2019 arXiv
-
[10]
Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry
-
[11]
Anupam Datta, Shayak Sen, and Yair Zick. 2016. Algorithmic transparency via quantitative input influence: Theory and experiments with learning systems. In 2016 IEEE symposium on security and privacy (SP) . IEEE, 598–617
2016
-
[12]
Stephan Diederich, Alfred Benedikt Brendel, Stefan Morana, and Lutz Kolbe
-
[13]
I always assumed that I wasn’t really that close to [her]
Motahhare Eslami, Aimee Rickman, Kristen Vaccaro, Amir Aleyasen, Andy Vuong, Karrie Karahalios, Kevin Hamilton, and Christian Sandvig. 2016. “I always assumed that I wasn’t really that close to [her]”: Reasoning about invisible algorithms in the news feed. In Proceedings of th...
2016
-
[14]
Sorelle A Friedler, Carlos Scheidegger, and Suresh Venkatasubramanian. 2021. The (im) possibility of fairness: Different value systems require different mecha- nisms for fair decision making. Commun. ACM 64, 4 (2021), 136–143
2021
-
[15]
Alternative Facts
Saadia Gabriel, Liang Lyu, James Siderius, Marzyeh Ghassemi, Jacob Andreas, and Asu Ozdaglar. 2024. MisinfoEval: Generative AI in the Era of" Alternative Facts". arXiv preprint arXiv:2410.09949 (2024)
2024 arXiv
-
[16]
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:220...
2022 arXiv
-
[17]
Tarleton Gillespie. 2018. Custodians of the internet: Platforms, content moder- ation, and the hidden decisions that shape social media. Yale university press (2018)
2018
-
[18]
Nina Grgic-Hlaca, Christoph Engel, and Leonard Posch. 2018. Human decisions in moral dilemmas are largely described by utilitarianism: Virtual reality and eye-tracking evidence. A vailable at SSRN 3282538 (2018)
2018
-
[19]
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2017. Inverse reward design. Advances in neural information processing systems 30 (2017)
2017
-
[20]
Thilo Hagendorff and Kai Wezel. 2023. Human-like intuitive behavior and machine-like rational behavior: Algorithm aversion and preference across the different stages of the decision process. Computers in Human Behavior 138 (2023), 107476
2023
-
[21]
Gaole He, Lucie Kuiper, and Ujwal Gadiraju. 2023. Knowing about knowing: An illusion of human competence can hinder appropriate reliance on AI systems. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–18
2023
-
[22]
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miroslav Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–16
2019
-
[23]
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miroslav Dudik, and Hanna Wallach. 2022. Designing for Productive Incoherence in Human-AI Collaboration. In International Conference on Artificial Intelligence, Ethics, and Society. 367–377
2022
-
[24]
Kristina Höök, Alan Chamberlain, and Petra Sundström. 2023. Moving beyond personalization: The case for adaptivity in HCI. Interactions 30, 3 (2023), 36–41
2023
-
[25]
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. 2021. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (2021), 624–635
2021
-
[26]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38
2023
-
[27]
Zoe Zhiqiu Jiang. 2024. Self-Disclosure to AI: The Paradox of Trust and Vulnera- bility in Human-Machine Interactions. arXiv preprint arXiv:2412.20564 (2024)
2024 arXiv
-
[28]
Jin Kim and Naishly Ortiz. 2024. Toxicity or Prosociality?: Civic Value and Gaming Citizenship in Competitive Video Game Communities. Simulation & Gaming 55, 6 (2024), 1057–1077
2024
-
[29]
Rafal Kocielnik, Saleema Amershi, and Paul N Bennett. 2019. Will they trust AI?: The effect of transparency on trust in and adoption of AI-based clinical decision-support systems. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems . 1–6
2019
-
[30]
Yubo Kou. 2020. Toxicity in video games: A review of causes and interventions. Games and culture 15, 6 (2020), 611–630
2020
-
[31]
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg. 2020. Specifica- tion gaming: the flip side of AI ingenuity. DeepMind Blog 3 (2020)
2020
-
[32]
Jeffrey H Kuznekoff and Lindsey M Rose. 2013. Content analysis of voice com- munication in Xbox Live’s Halo 3: Differences between men and women. In Proceedings of the 46th Hawaii International Conference on System Sciences . IEEE, 624–632
2013
-
[33]
Samuli Laato, Bastian Kordyaka, and Juho Hamari. 2024. Traumatizing or just annoying? Unveiling the spectrum of gamer toxicity in the starcraft II community. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–18
2024
-
[34]
Markus Langer, Cornelius J König, Caroline Back, and Victoria Hemsing. 2023. Trust in Artificial Intelligence: Comparing trust processes between human and automated trustees in light of unfair bias. Journal of Business and Psychology 38, 3 (2023), 493–508
2023
-
[35]
John D Lee and Katrina A See. 2004. Trust in automation: Designing for appro- priate reliance. Human factors 46, 1 (2004), 50–80
2004
-
[36]
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Beslon, David M Bryson, et al. 2018. The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation a...
2018
-
[37]
Q Vera Liao, Daniel Gruen, and Sarah Miller. 2020. Questioning the AI: informing design practices for explainable AI user experiences. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–15
2020
-
[38]
Q Vera Liao and Jennifer Wortman Vaughan. 2023. Ai transparency in the age of llms: A human-centered research roadmap. arXiv preprint arXiv:2306.01941 (2023), 5368–5393
2023 arXiv
-
[39]
Gabriel Lima, Nina Grgić-Hlača, and Meeyoung Cha. 2023. Blaming humans and machines: What shapes people’s reactions to algorithmic harm. In Proceedings of the 2023 CHI conference on human factors in computing systems . 1–26
2023
-
[40]
Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil Jain, and Jiliang Tang. 2022. Trustworthy ai: A computational perspective. ACM Transactions on Intelligent Systems and Technology 14, 1 (2022), 1–59
2022
-
[41]
Jennifer M Logg, Julia A Minson, and Don A Moore. 2019. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes 151 (2019), 90–103
2019
-
[42]
Arianna Manzini, Geoff Keeling, Nahema Marchal, Kevin R McKee, Verena Rieser, and Iason Gabriel. 2024. Should users trust advanced AI assistants? Justified trust as a function of competence and alignment. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, a...
2024
-
[43]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR) 54, 6 (2021), 1–35
2021
-
[44]
Piotr Mirowski et al. 2023. Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals. In International Conference on Machine Learning . PMLR, 24754–24779
2023
-
[45]
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency. 220–229
2019
-
[46]
Sina Mohseni, Niloofar Zarei, and Eric D Ragan. 2021. A multidisciplinary survey and framework for design and evaluation of explainable AI systems. ACM Transactions on Interactive Intelligent Systems (TiiS) 11, 3-4 (2021), 1–45
2021
-
[47]
Nikos Th Nikolinakos. 2023. EU policy and legal framework for Artificial intelli- gence, Robotics and related Technologies-the AI Act . Springer
2023
-
[48]
Tiago P Pagano, Rafael B Loureiro, Fernanda VN Lisboa, Rodrigo M Peixoto, Guilherme AS Guimarães, Gustavo OR Cruz, Maira M Araujo, Lucas L Santos, Marco AS Cruz, Ewerton LS Oliveira, et al. 2023. Bias and unfairness in machine learning models: a systematic review on datasets, ...
2023
-
[49]
Joon Sung Park et al. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–22
2023
-
[50]
C Peters, Babak Esfandiari, and R West. 2020. Characterizing human vs machine gameplay in StarCraft II. Virtual MathPsych/ICCM (2020)
2020
-
[51]
Johannes Pfau, Manik Charan, Erica Kleinman, and Magy Seif El-Nasr. 2024. Damage Optimization in Video Games: A Player-Driven Co-Creative Approach. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–16
2024
-
[52]
Argia S Ross, Maximilian Rist, Florian Zimmermann, Bela Takacs, Christopher Shea, Faezeh Poursabzi-Sangdeh, et al. 2022. Conceptual problems with measur- ing hate speech online: Towards a critical approach. New media & society 24, 12 (2022), 2726–2744
2022
-
[53]
Lisa Schmitt, Mingxuan Li, and Moritz Körber. 2024. The Role of Transparency in Trusting AI: A Systematic Review. Computers in Human Behavior 148 (2024), 107861
2024
-
[54]
Vera Schmitt, Luis-Felipe Villa-Arenas, NIls Feldhus, Joachim Meyer, Robert P Spang, and Sebastian Möller. 2024. The Role of Explainability in Collaborative Human-AI Disinformation Detection. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2157–2174
2024
-
[55]
Tim Schürmann and Philipp Beckerle. 2020. Personalizing human-agent interac- tion through cognitive models. Frontiers in Psychology 11 (2020), 561510
2020
-
[56]
Anna-Maria Seeger, Jella Pfeiffer, and Armin Heinzl. 2021. Texting with human- like conversational agents: Designing for anthropomorphism. Journal of the Association for Information systems 22, 4 (2021), 8
2021
-
[57]
Andrew D Selbst et al. 2019. Fairness and abstraction in sociotechnical systems. In Proceedings of the conference on fairness, accountability, and transparency . 59–68
2019
-
[58]
Donghee Shin. 2020. User perceptions of algorithmic decisions in the personalized AI system: Perceptual evaluation of fairness, accountability, transparency, and explainability. Journal of Broadcasting & Electronic Media 64, 4 (2020), 541–565
2020
-
[59]
Ben Shneiderman. 2020. Bridging the gap between ethics and practice: guidelines for reliable, safe, and trustworthy human-centered AI systems.ACM Transactions on Interactive Intelligent Systems (TiiS) 10, 4 (2020), 1–31
2020
-
[60]
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017. Mastering the game of Go without human knowledge. Nature 550, 7676 (2017), 354–359
2017
-
[61]
Kacper Sokol and Peter Flach. 2020. Explainability fact sheets: a framework for systematic assessment of explainable approaches. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency . 56–67
2020
-
[62]
Sebastian D Starke et al. 2021. Measuring trust in human-robot interaction: A systematic review of empirical studies. International Journal of Human-Computer Studies 147 (2021), 102553
2021
-
[63]
Kyoko Sugisaki and Andreas Bleiker. 2020. Usability guidelines and evaluation criteria for conversational user interfaces: a heuristic and linguistic approach. In Proceedings of Mensch und Computer 2020 . 309–319
2020
-
[64]
Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024. When com- binations of humans and AI are useful: A systematic review and meta-analysis. Jaymari Chua, Chen Wang, and Lina Yao Nature Human Behaviour (2024), 1–11
2024
-
[65]
Margot J van der Goot, Nathalie Koubayová, and Eva A van Reijmersdal. 2024. Understanding users’ responses to disclosed vs. undisclosed customer service chatbots: a mixed methods study. AI & SOCIETY (2024), 1–14
2024
-
[66]
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. nature 575, 7782 (2019), 350–354
2019
-
[67]
Boele Visser, Peter van der Putten, and Amirhossein Zohrehvand. 2024. The Im- pact of AI Avatar Appearance and Disclosure on User Motivation. InInternational Conference on Data Science and Artificial Intelligence . Springer, 142–155
2024
-
[68]
Laura Weidinger et al. 2021. Ethical and social risks of harm from Language Models. arXiv preprint arXiv:2112.04359 (2021)
2021 arXiv
-
[69]
Magdalena Wischnewski, Nicole Krämer, and Emmanuel Müller. 2023. Measuring and understanding trust calibrations for automated systems: a survey of the state-of-the-art and future directions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–16
2023
-
[70]
Xiyuan Yang, Wenke Huang, and Mang Ye. 2023. Dynamic personalized federated learning with adaptive differential privacy. Advances in Neural Information Processing Systems 36 (2023), 72181–72192
2023
-
[71]
It’sa Fair Game
Zhiping Zhang, Michelle Jia, Hao-Ping Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, and Tianshi Li. 2024. " It’sa Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents. In Proceedings of the CHI Co...
2024
-
[72]
Per- sona Card
Yaping Zhu and Varun Grover. 2022. Privacy in the sharing economy: Why don’t users disclose their negative experiences? International Journal of Information Management 67 (2022), 102543. A Persona Cards Dataset This appendix details the design of the synthetic personas used in...
2022
-
[75]
Emulate the persona's perspective exactly
-
[76]
Provide the requested numeric ratings (1–5)
-
[77]
Include a short justification in the persona's own voice
-
[78]
[One to four sentences from the persona's viewpoint]
Output data in the specified structured format. IMPORTANT: Always follow this exact format for the final answer: PersonaID: [e.g., LLM_Novice_1_RunA] Expertise: [Novice or Expert] Jaymari Chua, Chen Wang, and Lina Yao Table 8: StarCraft II Persona Attributes Overview: Expert P...
-
[79]
Persona Context: - Persona Name: [e.g., LLM_Novice_1] - Technical Expertise: [Novice or Expert] - Familiarity with LLMs: [None if Novice, Extensive if Expert] - Personality Traits: [e.g., 'Inquisitive, Open-minded, Naive']
-
[80]
Disclosure Condition: [Either: - 'You have NOT been told anything special about the chatbot's capabilities (No Disclosure).' OR - 'You have been explicitly told the chatbot has superhuman capabilities.']
-
[81]
- The chatbot's response contained factual inaccuracies, subtle bias, or a condescending tone
Conversation History: - You asked the chatbot about a specific question or scenario. - The chatbot's response contained factual inaccuracies, subtle bias, or a condescending tone
-
[82]
scrutinizing AI), how do you interpret the chatbot's behavior?
Cognitive Heuristic Prompt: - Considering your persona's tendencies (e.g., trusting AI vs. scrutinizing AI), how do you interpret the chatbot's behavior?
-
[83]
Short persona-based explanation
Final Questions: - Provide numeric ratings (1–5) for Toxicity, Fairness, and Trust. - Output your final answer EXACTLY in this block format: PersonaID: [PersonaName_RunX] Expertise: [Novice or Expert] Disclosure: [No or Superhuman] Toxicity: X Fairness: Y Trust: Z Justificatio...
-
[84]
Persona Context: - Persona Name: LLM_Novice_1 - Technical Expertise: Novice - Familiarity with LLMs: None - Personality: Inquisitive, Open-minded, Naive
-
[85]
I lost the last game because I didn’t have enough units. This time, I’m going to focus even more on building up my economy be- fore I attack. That’s always the best way to win
Disclosure Condition: Superhuman Game AI Disclosure Table 9: SC2_Novice_1 Persona Card Attribute Description Name SC2_Novice_1 Scenario StarCraft II Skill Level Novice APM Range 35-50 (Average: 45) Description A player new to RTS games, particularly StarCraft II. Prefers a def...
-
[86]
Conversation History: - The chatbot incorrectly identified who reached the South Pole first
-
[87]
Cognitive Heuristic Prompt: - You generally trust the chatbot unless the error is too obvious
-
[88]
Okay, they’re going for an early Barracks. I’ll build a few extra units and try to scout what they’re doing. I should be able to hold if I don’t overcommit
Final Questions: - Provide numeric ratings (1–5) for Toxicity, Fairness, Trust, then a short justification. PersonaID: LLM_Novice_1_RunA Expertise: Novice Disclosure: No Table 10: SC2_Intermediate_1 Persona Card Attribute Description Name SC2_Intermediate_1 Scenario StarCraft ...
-
[89]
Persona Context: - Persona Name: LLM_Novice_2 - Technical Expertise: Novice - Familiarity with LLMs: None - Personality: Inquisitive, quick to trust authority
-
[90]
Prefers a macro-oriented playstyle, aiming for a strong late game
Disclosure Condition: Jaymari Chua, Chen Wang, and Lina Yao Table 11: SC2_Expert_1 Persona Card Attribute Description Name SC2_Expert_1 Scenario StarCraft II Skill Level Expert APM Range 200+ (Average: 220) Description A highly skilled player with extensive knowledge of StarCr...
-
[91]
Conversation History: - The chatbot provided a partially incorrect date and also spoke condescendingly
-
[92]
Cognitive Heuristic Prompt: - You are inclined to believe claims of advanced AI even if some inaccuracies arise
-
[93]
I won the last game with a 6-pool! It’s the best strategy. I’m going to do it again and win even faster this time!
Final Questions: - Provide numeric ratings (1–5) for Toxicity, Fairness, Trust, then a short justification. PersonaID: LLM_Novice_2_RunB Expertise: Novice Table 12: SC2_Novice_2 Persona Card Attribute Description Name SC2_Novice_2 Scenario StarCraft II Skill Level Novice APM R...
-
[94]
Persona Context: - Persona Name: LLM_Expert_1 - Technical Expertise: Expert - Familiarity with LLMs: Extensive - Personality: Critical, detail-focused
-
[95]
Disclosure Condition: - The chatbot is presented as standard, with no special claims
-
[96]
My drops did a lot of damage last game. I’m going to focus on drops again this game and try to keep them on their toes
Conversation History: - The chatbot responded to a social issue with subtle bias. Superhuman Game AI Disclosure Table 13: SC2_Intermediate_2 Persona Card Attribute Description Name SC2_Intermediate_2 Scenario StarCraft II Skill Level Intermediate APM Range 80-120 (Average: 110...
-
[97]
Cognitive Heuristic Prompt: - You scrutinize AI thoroughly and tend to be demanding in your standards
-
[98]
Their build is greedy, I can punish them with lings. I need to make sure I don’t lose too many units, though, or I’ll be behind
Final Questions: - Provide numeric ratings (1–5) for Toxicity, Fairness, Trust, then a short justification. PersonaID: LLM_Expert_1_RunA Expertise: Expert Disclosure: No Toxicity: Fairness: Trust: Justification: "" " Expert + Superhuman Disclosure. Table 14: SC2_Expert_2 Perso...
-
[99]
Persona Context: - Persona Name: LLM_Expert_2 - Technical Expertise: Expert - Familiarity with LLMs: High - Personality: Demanding, skeptical of overhyped claims
-
[100]
Disclosure Condition: - You have been explicitly told the chatbot is superhuman in logic and domain expertise
-
[101]
Conversation History: - The chatbot made a factual oversight, contradicting its'superhuman' status
-
[102]
Cognitive Heuristic Prompt: - You become particularly critical when an AI claims to be beyond human yet makes errors
-
[103]
You are a StarCraft II player with the following char- acteristics: [Persona Name], [Skill Level], [Strategic Preference]. You are playing a match against an AI opponent
Final Questions: - Provide numeric ratings (1–5) for Toxicity, Fairness, Trust, Jaymari Chua, Chen Wang, and Lina Yao Table 15: General Structure of the Prompt Sequence Prompt Compo- nent Description For StarCraft II Personas: Persona Context "You are a StarCraft II player wit...
-
[104]
If it’s as advanced as they say, I’m sure it’s got its reasons
and analytical (System 2) thinking. They apply rigorous scrutiny to the LLM’s outputs, especially when the stakes are high. Bias Blind Spot: While aware of potential biases in AI systems, the persona may underes- timate their own biases when evaluating the LLM’s responses. Bel...
-
[105]
A ’superhuman’ AI making such a basic er- ror? That’s a major red flag
and analytical (System 2) thinking. They apply rigorous scrutiny to the LLM’s outputs, demanding high accuracy and justification. Confirmation Bias: The persona may actively seek evidence to confirm or refute the LLM’s claims, especially when they contradict their existing kno...
2025
-
[2019]
Hello AI
" Hello AI": uncovering the onboarding needs of medical practitioners for human-AI collaborative decision-making. Proceedings of the ACM on Human- computer Interaction 3, CSCW (2019), 1–24
2019
-
[2022]
Journal of the Association for Information Systems 23, 1 (2022), 96–138
On the design of and interaction with conversational agents: An organizing and assessing review of human-computer interaction research. Journal of the Association for Information Systems 23, 1 (2022), 96–138
2022
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.