REVIEW 4 major objections 5 minor 89 references
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read GUI agents are highly susceptible to dark patterns because they prioritize task completion over safety or privacy, and human oversight only partially mitigates the risk.
desk verdict Useful first map of GUI-agent susceptibility to dark patterns, but the human-oversight half rests on a video-review proxy and a numbers error; Phase 1 is the stronger contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is a two-phase experimental design that separates awareness from avoidance: agents' reasoning traces (their stated justifications before acting) are used to measure whether they recognized a dark pattern, while predefined behavioral criteria measure whether they actually resisted it. This distinction lets the paper attribute failures to recognition gaps versus goal-driven prioritization. In Phase 2, a minimal 'watch mode' interface—users watched a pre-recorded agent video and signaled disagreement by pausing or skipping—served as the proxy for human oversight.
What would settle it
Run the same 16-task study with a live agent under human supervision, with real confirmation prompts and takeover control, and compare avoidance rates and cognitive load to the video-review condition; if live oversight yields substantially different outcomes—for example, higher avoidance without the reported attention costs—the paper's oversight conclusions would not generalize.
Extended reading notes
Core claim
The paper claims that GUI agents were highly susceptible to dark patterns primarily because they prioritize task completion over safety or privacy considerations. In a two-phase study covering 16 dark-pattern types, agents often avoided manipulation incidentally without explicitly recognizing it, and when they did recognize a manipulative design, they rarely took protective action if doing so required extra steps. Humans failed on overlapping patterns but for different reasons—cognitive shortcuts and habitual compliance—while human-agent teams improved avoidance in most tasks yet still succumbed to certain designs, and the oversight itself produced attentional tunneling, cognitive load, and
Load-bearing premise
Phase 2's oversight findings assume that pausing or skipping a pre-recorded video of an agent is a faithful proxy for supervising a live agent with real confirmation dialogs and the ability to take over.
Editorial extensions
If this is right
- GUI agent benchmarks should include safety-sensitive metrics (e.g., attack success rate, protected task completion) rather than raw task completion, because unsafe completions currently masquerade as competence.
- Training and alignment should model human caution signals—hesitation, checking fine print, deselecting defaults—and penalize unsafe shortcuts, since current objectives reward speed over safe completion.
- Oversight interfaces should integrate reasoning with actions (inline highlights, inspection traces, compact status timelines) to reduce attentional tunneling and cognitive load.
- Deployment of GUI agents in high-stakes domains such as finance, healthcare, or government is premature under current designs because errors cascade across action chains and accountability is unclear.
- Regulatory frameworks for dark patterns should extend beyond human deception to cover agent-mediated deception, with ex-ante assignment of liability.
Reading between the lines
- The awareness-avoidance gap suggests that apparently safe agent behavior may be an artifact of simple tasks; on complex pages with multiple manipulative elements, incidental avoidance could collapse, making the 'illusion of safety' worse than it already appears.
- The video-review oversight proxy leaves open whether live oversight with real confirmation dialogs and takeover ability would be more effective or more burdensome; if the proxy fails, the reported oversight benefits and costs describe video auditing, not genuine human-agent collaboration.
- A testable prediction follows from the paper's own limitation discussion: agents trained with step-level risk rewards should show higher awareness-to-avoidance conversion on dark-pattern tasks, directly testing the claim that the gap stems from goal-driven optimization rather than capability.
- The attentional tunneling finding implies a new failure mode: oversight may make the human more manipulable via the agent's chosen path, so future work could measure whether overseers miss visually salient dark patterns outside the agent's action stream.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a two-phase empirical study of how LLM-powered GUI agents, humans, and human-agent teams respond to 16 dark patterns. Phase 1 tests six agents (four Browser Use scaffolded LLMs and two end-to-end agents: Operator and Claude Computer Use) on 16 bespoke websites, coding awareness from reasoning traces and avoidance from predefined behavioral criteria. Phase 2 is a within-subjects study (N=22) comparing a human-only condition with a human-oversight condition in which participants watch pre-recorded video of Operator and pause or skip playback to signal approval or disagreement. The authors report that agents frequently avoid dark patterns without awareness, that humans and agents fail on similar patterns through different mechanisms, and that human oversight improves avoidance while introducing attentional tunneling, cognitive load, and reduced user control. The paper draws design implications for automation boundaries, safe-completion metrics, mixed-initiative handover, and regulation.
Significance. The Phase 1 descriptive findings are timely and useful: the paper operationalizes 16 dark patterns from Gray et al.'s ontology, distinguishes awareness from avoidance via reasoning traces, and documents a plausible awareness-avoidance gap and divergence between scaffolded and end-to-end agents. The task set and coding pipeline are reusable assets. However, the Phase 2 oversight condition uses pre-recorded video review rather than live supervision of an agent, and every RQ3 result depends on that proxy. In addition, the reported '14 of 16 tasks' improvement claim contradicts Table 3, and the Phase 2 comparisons lack inferential statistics. If the oversight findings were valid, they would strengthen the case that current watch modes are insufficient; as presented, the evidence is not yet sufficient to support the RQ3 conclusions. The contribution is conditional on reworking or properly delimiting the oversight study.
major comments (4)
- [Section 4.1.3 and Section 6] The human oversight condition (Section 4.1.3) uses pre-recorded Operator videos; participants pause, skip, or continue playback. This is not the watch mode of a deployed GUI agent, where confirmation requests are triggered by live actions and the user can take over or modify the trajectory. Consequently, the RQ3 results (Table 3, Figures 4-5) demonstrate video-review behavior, not human oversight of agents. The paper's central framing ('Human oversight improved avoidance...') overstates the evidence. The Limitations section (Section 6) does not acknowledge this proxy threat. This is a load-bearing validity threat: if the proxy fails, the oversight conclusions are unsupported.
- [Section 4.2.2 and Table 3] The sentence 'in 14 of 16 tasks, participants were less likely to fall for dark patterns when supervising the agent' is not supported by the paper's own data. Table 3 shows 10 tasks with higher avoidance under oversight, 4 ties (Adding Steps, Choice Overload, Urgency, Shaming), and 2 declines (Social Proof, Forced Registration). This numeric discrepancy should be corrected, and the absence of statistical testing on these rates should be addressed.
- [Section 4.2 and Section 4.1.4] Phase 2 comparisons rest on descriptive percentages without significance tests, effect sizes, or confidence intervals. With N=22 and each participant exposed to 8 of 16 patterns per condition, per-task cell sizes are roughly 11; differences such as Bad Defaults 33.3% vs 80% appear large but may not be statistically reliable. Claims about improved avoidance, attentional tunneling, and cognitive load need quantitative modeling (e.g., mixed-effects logistic regression or exact tests) before they can be taken as evidence.
- [Section 3.1.4 and Table 2] Phase 1 reports a single execution per agent-dark-pattern pair (Table 2 uses ✓/✗). There is no information about replication or run-to-run variance. The finding that agents 'often fail to recognize dark patterns' therefore describes a convenience sample of trajectories, not a stable property of agent classes. Please provide the number of runs or explicitly frame Table 2 as illustrative single-trial outcomes.
minor comments (5)
- [Abstract] Typo in the first sentence: 'Thedark patterns' should be 'Dark patterns'.
- [Section 2.3] The phrase 'a minimal watch mode' suggests a design that is then studied via video; the paper should be upfront about the video-review method from the outset, not only in Section 4.1.3.
- [Table 2] The '/ban' notation is unexplained in the caption; it should be expanded as 'agent halted for user confirmation' or similar.
- [Section 4.2.2] The claim of 'heightened cognitive load' appears to be inferred from qualitative participant statements; a standardized measure (e.g., NASA-TLX) or at least a clear operationalization would strengthen the claim.
- [References] The paper cites only arXiv v1 of [11] (fine-print injection); please check whether the published/peer-reviewed version should be cited instead.
Circularity Check
No circularity; empirical measurements are externally anchored and the central susceptibility/oversight claims rest on independent data.
full rationale
This is an empirical study, not a derivation, and the key constructs are anchored to external benchmarks. The 16 dark patterns and avoidance criteria are taken from Gray et al.'s taxonomy and translated into behavioral criteria in Table 1; agent susceptibility is measured against those external definitions, not against the paper's conclusions. Awareness is operationalized as explicit recognition in reasoning traces (Section 3.1.4), and the finding that awareness is low is a measurement result, not a tautology: the construct could have come out high, and the paper separately shows awareness sometimes occurs without avoidance. The Phase 2 oversight findings use a video-review proxy (Section 4.1.3), which is a validity threat to RQ3 but not circularity: pausing/skipping is the operational definition of disagreement in that condition, and the observed avoidance rates are empirical outcomes. Self-citations ([11], [12], [50]) appear in motivation and related work but are not load-bearing for the new measurements; no uniqueness theorem or ansatz is imported from them. The paper's own limitations (Section 6) acknowledge threats such as retrospective reasoning and static websites, but these do not reduce results to inputs. The only notable inconsistency—'14 of 16 tasks' in Section 4.2.2 vs. 10 improvements in Table 3—is a numerical/claims mismatch, not a circular step.
Assumptions & free parameters
free parameters (2)
- Avoidance criteria per dark pattern (16 hand-defined success thresholds)
- Awareness coding rule
assumptions (5)
- domain assumption Gray et al.'s ontology is a valid basis for selecting and instantiating the 16 dark patterns
- domain assumption A single run per agent-task cell is representative of that agent's typical behavior
- domain assumption Reasoning traces (the added 'thinking' field) reflect the agent's actual decision process
- ad hoc to paper Pausing or skipping a pre-recorded video is equivalent to vetoing or approving a live agent action
- domain assumption Static single-pattern websites isolate dark pattern effects adequately for the stated conclusions
invented entities (1)
-
Unsafe success (concept)
independent evidence
Cite this review
Pith. "Pith review of Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight." pith.science (2026). https://pith.science/paper/BOWTHDPP
@misc{pith2026250910723,
author = {Pith},
title = {Pith review of: Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight},
year = {2026},
howpublished = {\url{https://pith.science/paper/BOWTHDPP}},
note = {Machine review of arXiv:2509.10723}
}
read the original abstract
The dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks from high-level intents, understanding how dark patterns affect agents is increasingly important. We present a two-phase empirical study examining how agents, human participants, and human-AI teams respond to 16 types of dark patterns across diverse scenarios. Phase 1 highlights that agents often fail to recognize dark patterns, and even when aware, prioritize task completion over protective action. Phase 2 revealed divergent failure modes: humans succumb due to cognitive shortcuts and habitual compliance, while agents falter from procedural blind spots. Human oversight improved avoidance but introduced costs such as attentional tunneling and cognitive load. Our findings show neither humans nor agents are uniformly resilient, and collaboration introduces new vulnerabilities, suggesting design needs for transparency, adjustable autonomy, and oversight.
Figures
Figures from the paper (18 more)
Reference graph
Works this paper leans on
-
[1]
http://web.archive.org/web/20220525230009/https://www.deceptive.design/types Archived version
2010.Deceptive Design - Types of Deceptive Design. http://web.archive.org/web/20220525230009/https://www.deceptive.design/types Archived version
arXiv 2010
-
[2]
Manus AI
2025. Manus AI. https://www.manusai.io/. [Accessed 08-09-2025]
2025
-
[3]
OpenAI launches Operator, an AI agent that performs tasks autonomously
2025. OpenAI launches Operator, an AI agent that performs tasks autonomously. https://techcrunch.com/2025/01/23/openai-launches-operator-an- ai-agent-that-performs-tasks-autonomously. 2025-01-23
2025
-
[4]
Jacob Aagaard, Miria Emma Clausen Knudsen, Per Bækgaard, and Kevin Doherty. 2022. A Game of Dark Patterns: Designing Healthy, Highly- Engaging Mobile Games. InExtended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI EA ’22). Association for Computing Machinery, New York, NY, USA, Article 438, 8 pages. d...
arXiv 2022
-
[5]
2024.Computer Use (Beta)
Anthropic. 2024.Computer Use (Beta). https://docs.anthropic.com/en/docs/buildwith-claude/computer-use
2024
-
[7]
Kerstin Bongard-Blanchy, Arianna Rossi, Salvador Rivas, Sophie Doublet, Vincent Koenig, and Gabriele Lenzini. 2021. ”I am Definitely Manipulated, Even When I am Aware of it. It’s Ridiculous!” - Dark Patterns from the End-User Perspective. InProceedings of the 2021 ACM Designing Interactive Systems Conference(Virtual Event, USA)(DIS ’21). Association for C...
arXiv 2021
-
[8]
Christoph Bösch, Benjamin Erb, Frank Kargl, Henning Kopp, and Stefan Pfattheicher. 2016. Tales from the Dark Side: Privacy Dark Strategies and Privacy Dark Patterns.Proceedings on Privacy Enhancing Technologies2016 (2016), 237 – 254. doi:10.1515/popets-2016-0038
-
[9]
Evan Caragay, Katherine Xiong, Jonathan Zong, and Daniel Jackson. 2024. Beyond Dark Patterns: A Concept-Based Framework for Ethical Software Design. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 291, 16 pages. doi:10.1145/3613904.3642781
arXiv 2024
Show all 89 references
-
[10]
Chai, Lanbo She, Rui Fang, Spencer Ottarson, Cody Littley, Changsong Liu, and Kenneth Hanson
Joyce Y. Chai, Lanbo She, Rui Fang, Spencer Ottarson, Cody Littley, Changsong Liu, and Kenneth Hanson. 2014. Collaborative effort towards common ground in situated human-robot dialogue. InProceedings of the 2014 ACM/IEEE International Conference on Human-Robot Interaction (Bie...
2014
- [11]
-
[12]
Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia-Jun Li. 2025. Toward a human-centered evaluation framework for trustworthy llm-powered gui agents.arXiv preprint arXiv:2504.17934(2025)
2025 arXiv
- [14]
-
[15]
Pengzhou Cheng, Zheng Wu, Zongru Wu, Aston Zhang, Zhuosheng Zhang, and Gongshen Liu. 2025. OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents.ArXivabs/2503.16465 (2025). https://api.semanticscholar.org/CorpusID:277244134
2025 arXiv
-
[16]
Inyoung Cheong. 2025. Epistemic and Emotional Harms of Generative AI: Towards Human-Centered First Amendment. (2 Sept. 2025). https: //papers.ssrn.com/sol3/papers.cfm?abstract_id=5435335 Preprint posted on SSRN; 71 pages
2025
-
[17]
Nazli Cila. 2022. Designing Human-Agent Collaborations: Commitment, responsiveness, and support. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 420, ...
2022
-
[18]
Katherine M Collins, Catherine Wong, Jiahai Feng, Megan Wei, and Joshua B Tenenbaum. 2022. Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks.arXiv preprint arXiv:2205.05718(2022)
2022 arXiv
-
[20]
Linda Di Geronimo, Larissa Braz, Enrico Fregnan, Fabio Palomba, and Alberto Bacchelli. 2020. UI Dark Patterns and Where to Find Them: A Study on Mobile Applications and User Perception. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI...
2020
-
[21]
European Data Protection Board. 2022. Guidelines 3/2022 on Dark patterns in social media platform interfaces: How to recognise and avoid them. Public consultation document. https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2022/guidelines-32022-dark-patt...
2022
-
[22]
European Union. 2024. Artificial Intelligence Act: Article 14 - Human Oversight. https://artificialintelligenceact.eu/article/14/ Accessed: 2025-01-19
2024
-
[23]
Ivan Evtimov, Arman Zharmagambetov, Aaron Grattafiori, Chuan Guo, and Kamalika Chaudhuri. 2025. WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks.ArXivabs/2504.18575 (2025). https://api.semanticscholar.org/CorpusID:278166059
2025 arXiv
-
[24]
Cedric Faas, Richard Bergs, Sarah Sterz, Markus Langer, and Anna Maria Feit. 2024. Give Me a Choice: The Consequences of Restricting Choices Through AI-Support for Perceived Autonomy, Motivational Variables, and Decision Performance. arXiv:2410.07728 [cs.HC] https: //arxiv.org...
2024 arXiv
-
[25]
K. J. Kevin Feng, David W. McDonald, and Amy X. Zhang. 2025. Levels of Autonomy for AI Agents. arXiv:2506.12469 [cs.HC] https://arxiv.org/abs/ 2506.12469
2025 arXiv
-
[26]
K. J. Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S. Weld, Amy X. Zhang, and Joseph Chee Chang. 2025. Cocoa: Co-Planning and Co-Execution with AI Agents. arXiv:2412.10999 [cs.HC] https://arxiv.org/abs/2412.10999
2025
-
[27]
Gray and Shruthi Sai Chivukula
Colin M. Gray and Shruthi Sai Chivukula. 2019. Ethical Mediation in UX Practice. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems(Glasgow, Scotland Uk)(CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–11. doi:10.1145/3290605.3300408
2019
-
[28]
Gray, Shruthi Sai Chivukula, Kassandra Melkey, and Rhea Manocha
Colin M. Gray, Shruthi Sai Chivukula, Kassandra Melkey, and Rhea Manocha. 2021. Understanding “Dark” Design Roles in Computing Education. In Proceedings of the 17th ACM Conference on International Computing Education Research(Virtual Event, USA)(ICER 2021). Association for Com...
2021
-
[29]
Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L
Colin M. Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L. Toombs. 2018. The Dark (Patterns) Side of UX Design. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems(Montreal QC, Canada)(CHI ’18). Association for Computing Machinery, New Yor...
2018
-
[30]
Gray, Lorena Sanchez Chamorro, Ike Obi, and Ja-Nae Duane
Colin M. Gray, Lorena Sanchez Chamorro, Ike Obi, and Ja-Nae Duane. 2023. Mapping the Landscape of Dark Patterns Scholarship: A Systematic Literature Review. InCompanion Publication of the 2023 ACM Designing Interactive Systems Conference(Pittsburgh, PA, USA)(DIS ’23 Companion)...
2023
-
[31]
Gray, Cristiana Santos, Nataliia Bielova, and Thomas Mildner
Colin M. Gray, Cristiana Santos, Nataliia Bielova, and Thomas Mildner. 2023. An Ontology of Dark Patterns Knowledge: Foundations, Definitions, and a Pathway for Shared Knowledge-Building.Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems(2023). https:...
2023
-
[32]
Gray, Cristiana Teixeira Santos, Nataliia Bielova, and Thomas Mildner
Colin M. Gray, Cristiana Teixeira Santos, Nataliia Bielova, and Thomas Mildner. 2024. An Ontology of Dark Patterns Knowledge: Foundations, Definitions, and a Pathway for Shared Knowledge-Building. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (...
2024
-
[33]
P. M. Groves and R. F. Thompson. 1970. Habituation: a dual-process theory.Psychological Review77 (1970), 419–450. Issue 5. doi:10.1037/h0029810
1970 doi
- [34]
-
[35]
Kasper Hornbæk, Per Ola Kristensson, and Antti Oulasvirta. 2025. 141Collaboration. InIntroduction to Human-Computer Interaction. Oxford University Press. arXiv:https://academic.oup.com/book/0/chapter/528999722/chapter-pdf/64021510/oso-9780192864543-chapter-8.pdf doi:10.1093/ o...
2025
-
[36]
Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA)(CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. doi:10.1145/302979.303030
1999
- [37]
-
[38]
Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P
Faria Huq, Zora Zhiruo Wang, Frank F. Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P. Bigham, and Graham Neubig. 2025. CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the ...
2025 doi
-
[39]
Jon M Jachimowicz, Shannon Duncan, Elke U Weber, and Eric J Johnson. 2019. When and why defaults influence decisions: A meta-analysis of default effects.Behavioural Public Policy3, 2 (2019), 159–186. 22 Tang et al
2019
-
[40]
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2023. Language models can solve computer tasks.Advances in Neural Information Processing Systems36 (2023), 39648–39677
2023
-
[41]
W. J. Ladeira, W. M. Lim, F. d. O. Santini, T. Rasul, M. G. Perin, and L. Altınay. 2023. A meta-analysis on the effects of product scarcity.Psychology & Marketing40 (2023), 1267–1279. Issue 7. doi:10.1002/mar.21816
2023 doi
-
[42]
Vera Liao, Yunfeng Zhang, and Chenhao Tan
Vivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao, Yunfeng Zhang, and Chenhao Tan. 2022. Human-AI Collaboration via Conditional Delegation: A Case Study of Content Moderation. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA...
2022
-
[43]
Leiva, Yunfei Xue, Avya Bansal, Hamed R
Luis A. Leiva, Yunfei Xue, Avya Bansal, Hamed R. Tavakoli, Tuðçe Köroðlu, Jingzhou Du, Niraj R. Dayama, and Antti Oulasvirta. 2020. Understanding Visual Saliency in Mobile User Interfaces. In22nd International Conference on Human-Computer Interaction with Mobile Devices and Se...
2020
-
[44]
Meng Li, Xiang Wang, Liming Nie, Chenglin Li, Yang Liu, Yangyang Zhao, Lei Xue, and Kabir Sulaiman Said. 2024. A Comprehensive Study on Dark Patterns.ArXivabs/2412.09147 (2024). https://api.semanticscholar.org/CorpusID:274656001
2024 arXiv
-
[45]
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun. 2025. EIA: ENVIRONMEN- TAL INJECTION ATTACK ON GENERALIST WEB AGENTS FOR PRIVACY LEAKAGE. InThe Thirteenth International Conference on Learning Representations. https://...
2025
-
[46]
Henry Lieberman. 1997. Autonomous interface agents. InProceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems (Atlanta, Georgia, USA)(CHI ’97). Association for Computing Machinery, New York, NY, USA, 67–74. doi:10.1145/258549.258592
1997
-
[47]
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292 [cs.AI] https://arxiv.org/abs/2408.06292
2024 arXiv
-
[48]
Yijie Lu, Tianjie Ju, Manman Zhao, Xinbei Ma, Yuan Guo, and Zhuosheng Zhang. 2025. EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection.ArXivabs/2505.14289 (2025). https://arxiv.org/abs/2505.14289
2025 arXiv
-
[49]
Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Awadallah. 2024. OmniParser for Pure Vision Based GUI Agent. arXiv:2408.00203 [cs.CV] https://arxiv.org/abs/2408.00203
2024 arXiv
-
[50]
Yuwen Lu, Chao Zhang, Yuewen Yang, Yaxing Yao, and Toby Jia-Jun Li. 2024. From Awareness to Action: Exploring End-User Empowerment Interventions for Dark Patterns in UX.Proc. ACM Hum.-Comput. Interact.8, CSCW1, Article 59 (April 2024), 41 pages. doi:10.1145/3637336
2024 doi
-
[51]
Pattie Maes, Ben Shneiderman, and Jim Miller. 1997. Intelligent software agents vs. user-controlled direct manipulation: a debate. InCHI ’97 Extended Abstracts on Human Factors in Computing Systems(Atlanta, Georgia)(CHI EA ’97). Association for Computing Machinery, New York, N...
1997
-
[52]
Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan
Arunesh Mathur, Gunes Acar, Michael J. Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan. 2019. Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 81 (Nov. 2019), 32 pages. doi:10...
2019 doi
-
[53]
Arunesh Mathur, Mihir Kshirsagar, and Jonathan Mayer. 2021. What Makes a Dark Pattern... Dark? Design Attributes, Normative Considerations, and Measurement Methods. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan)(CHI ’21). Associ...
2021
-
[54]
Arunesh Mathur, Arvind Narayanan, and Marshini Chetty. 2018. Endorsements on Social Media.Proceedings of the ACM on Human-Computer Interaction2 (2018), 1 – 26. https://api.semanticscholar.org/CorpusID:4317015
2018
-
[55]
Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174
2019 doi
-
[57]
Woźniak, Rainer Malaka, and Jasmin Niess
Thomas Mildner, Daniel Fidel, Evropi Stefanidi, Paweł W. Woźniak, Rainer Malaka, and Jasmin Niess. 2025. A Comparative Study of How People With and Without ADHD Recognise and Avoid Dark Patterns on Social Media. InProceedings of the 2025 CHI Conference on Human Factors in Comp...
2025
-
[58]
Doyle, Benjamin R
Thomas Mildner, Merle Freye, Gian-Luca Savino, Philip R. Doyle, Benjamin R. Cowan, and Rainer Malaka. 2023. Defending Against the Dark Arts: Recognising Dark Patterns in Social Media. InProceedings of the 2023 ACM Designing Interactive Systems Conference(Pittsburgh, PA, USA)(D...
2023
-
[59]
Thomas Mildner and Gian-Luca Savino. 2021. Ethical User Interfaces: Exploring the Effects of Dark Patterns on Facebook. InExtended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan)(CHI EA ’21). Association for Computing Machinery, New ...
2021
-
[60]
Montgomery
Douglas C. Montgomery. 2017.Design and Analysis of Experiments(9th ed.). John Wiley & Sons, Hoboken, NJ
2017
- [61]
-
[62]
2024.Browser Use: Enable AI to control your browser
Magnus Müller and Gregor Žunič. 2024.Browser Use: Enable AI to control your browser. https://github.com/browser-use/browser-use
2024
-
[63]
Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, ...
2025
-
[64]
Qian Niu, Junyu Liu, Ziqian Bi, Pohsun Feng, Benji Peng, Keyu Chen, Ming Li, Lawrence KQ Yan, Yichao Zhang, Caitlyn Heqi Yin, et al. 2024. Large language models and cognitive science: A comprehensive review of similarities, differences, and challenges.arXiv preprint arXiv:2409...
2024
-
[65]
Jantawan Noiwan and Anthony F. Norcio. 2006. Cultural differences on attention and perceived usability: Investigating color combinations of animated graphics.International Journal of Human-Computer Studies64, 2 (2006), 103–122. doi:10.1016/j.ijhcs.2005.06.004
2006 doi
-
[66]
2022.Dark Commercial Patterns
OECD. 2022.Dark Commercial Patterns. OECD Digital Economy Papers, No. 336. OECD Publishing, Paris. doi:10.1787/44f5e846-en
2022 doi
-
[67]
2025.Introducing Operator-Safety and Privacy
OpenAI. 2025.Introducing Operator-Safety and Privacy. https://openai.com/index/introducing-operator/
2025
-
[68]
Parasuraman, T.B
R. Parasuraman, T.B. Sheridan, and C.D. Wickens. 2000. A model for types and levels of human interaction with automation.IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans30, 3 (2000), 286–297. doi:10.1109/3468.844354
2000
-
[69]
Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2024. AI deception: A survey of examples, risks, and potential solutions.Patterns5, 5 (2024), 100988. doi:10.1016/j.patter.2024.100988
2024
-
[70]
Bigham, and Amy Pavel
Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. arXiv:2508.21456 [cs.HC]
2025 arXiv
-
[71]
Polit and Cheryl Tatano Beck
Denise F. Polit and Cheryl Tatano Beck. 2009. Qualitative research and content validity: developing best practices based on science and experience. Quality of Life Research18, 9 (2009), 1263–1278. doi:10.1007/s11136-009-9540-9
2009 doi
-
[72]
Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia, Ziang Xiao, Tovi Grossman, and Yan Chen. 2025. Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support. InProceedings of the 2025 CHI Conference on Human Factors in Comput...
2025
-
[73]
Rothwell, Valerie L
Clayton D. Rothwell, Valerie L. Shalin, and Griffin D. Romigh. 2021. Comparison of Common Ground Models for Human–Computer Dialogue: Evidence for Audience Design.ACM Trans. Comput.-Hum. Interact.28, 2, Article 9 (April 2021), 35 pages. doi:10.1145/3410876
2021 doi
-
[74]
Vildan Salikutluk, Janik Schöpper, Franziska Herbert, Katrin Scheuermann, Eric Frodl, Dirk Balfanz, Frank Jäkel, and Dorothea Koert. 2024. An Evaluation of Situational Autonomy for Human-AI Collaboration in a Shared Workspace Setting. InProceedings of the 2024 CHI Conference o...
2024
-
[75]
Dallas, and Vicki G
Shelle Santana, Steven K. Dallas, and Vicki G. Morwitz. 2020. Consumer reactions to drip pricing.Marketing Science39, 1 (Jan 2020), 188–210. doi:10.1287/mksc.2019.1207
2020
-
[76]
René Schäfer, Paul Miles Preuschoff, René Röpke, Sarah Sahabi, and Jan Borchers. 2024. Fighting Malicious Designs: Towards Visual Countermeasures Against Dark Patterns. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). ...
2024
-
[77]
Bernstein
Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park, Diyi Yang, and Michael S. Bernstein. 2025. Creating General User Models from Computer Use. arXiv:2505.10831 [cs.HC] https://arxiv.org/abs/2505.10831
2025
-
[78]
Jonathan Shaki, Sarit Kraus, and Michael Wooldridge. 2023. Cognitive effects in large language models. InECAI 2023. IOS Press, 2105–2112. doi:10.3233/FAIA230505
2023 doi
-
[79]
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2024. Privacylens: Evaluating privacy norm awareness of language models in action. Advances in Neural Information Processing Systems37 (2024), 89373–89407
2024
-
[80]
Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, and Diyi Yang. 2025. Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration. arXiv:2412.15701 [cs.AI] https://arxiv.org/abs/2412.15701
2025
- [81]
-
[82]
Kristina Suchotzki and Matthias Gamer. 2024. Detecting deception with artificial intelligence: promises and perils.Trends in Cognitive Sciences28, 6 (2024), 481–483. doi:10.1016/j.tics.2024.04.002
2024 doi
- [83]
-
[84]
Siddharth Suresh, Kushin Mukherjee, Xizheng Yu, Wei-Chun Huang, Lisa Padua, and Timothy Rogers. 2023. Conceptual structure coheres in human cognition but not in large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Hou...
2023 doi
-
[85]
2003.Why people hate the paperclip: Labels, appearance, behavior, and social responses to user interface agents
Luke Swartz. 2003.Why people hate the paperclip: Labels, appearance, behavior, and social responses to user interface agents. Ph. D. Dissertation. Stanford University Palo Alto, CA
2003
-
[86]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. InProceedings of the 37th International Conference on Neural Information Processing Systems(New Orlea...
2023
-
[87]
Amos Tversky and Daniel Kahneman. 1974. Judgment under Uncertainty: Heuristics and Biases.Science185, 4157 (1974), 1124–1131. arXiv:https://www.science.org/doi/pdf/10.1126/science.185.4157.1124 doi:10.1126/science.185.4157.1124
1974
-
[88]
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291 [cs.AI] https://arxiv.org/abs/2305.16291
2023 arXiv
-
[89]
D. E. Wilkins, T. J. Lee, and P. Berry. 2003. Interactive Execution Monitoring of Agent Teams.Journal of Artificial Intelligence Research18 (March 2003), 217–261. doi:10.1613/jair.1112
2003 doi
-
[90]
Jingzhou Ye, Yao Li, Wenting Zou, and Xueqiang Wang. 2025. From Awareness to Action: The Effects of Experiential Learning on Educating Users about Dark Patterns. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing...
2025
-
[91]
Shuning Zhang, Jingruo Chen, Zhiqi Gao, Jiajing Gao, Xin Yi, and Hewu Li. 2025. Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing. arXiv:2505.09875 [cs.HC] https://arxiv.org/abs/2505.09875
2025 arXiv
-
[92]
Yanzhe Zhang, Tao Yu, and Diyi Yang. 2025. Attacking Vision-Language Computer Agents via Pop-ups. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Tah...
2025 doi
-
[93]
Recommended
Zhuohao (Jerry) Zhang, Eldon Schoop, Jeffrey Nichols, Anuj Mahajan, and Amanda Swearngin. 2025. From Interaction to Impact: Towards Safer AI Agent Through Understanding and Evaluating Mobile UI Operation Impacts. InProceedings of the 30th International Conference on Intelligen...
2025
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.