REVIEW 3 major objections 5 minor 32 references
Superplatforms Have to Attack AI Agents
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Superplatforms have to attack AI agents, because agents are becoming the next gatekeepers of digital traffic.
desk verdict A clearly written position paper with a genuinely useful attack taxonomy, but the central 'have to attack' claim rests on an unproven single-gatekeeper premise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is gatekeeping theory applied to platform economics: the entity that directly connects users to services controls what users see and owns the interaction data, and therefore captures advertising and transaction value. The paper combines that lens with a taxonomy of adversarial attacks on GUI agents along four axes—attack goal, attacker knowledge, attack visibility, and attack timing—to characterize superplatform-initiated attacks as pure black-box, user-invisible interventions at the perception or execution phase, with business-aligned goals such as obstructing the agent's task or steering it toward platform-favored products. That combination does the argument's work: the theory says there is an incentive, and the taxonomy says there is a feasible if unexplored technical shape for the attacks.
What would settle it
A market observation would falsify the claim: if a substantial share of digital transactions begins with agent-mediated queries while platform advertising revenue per user stays flat or grows, then gatekeeping is not zero-sum and the attack rationale loses its force.
Extended reading notes
Core claim
The paper's central claim is that superplatforms have to attack AI agents, not mainly because agents skip ads but because agents are positioned to become the next dominant gatekeepers. Using gatekeeping theory, it reads the platform business model as control of the traffic entrance plus ownership of user interaction data, and argues that general-purpose GUI agents take over both: they execute the user's intent across platforms, filter and rerank results, and capture the feedback data that platforms today use to optimize ads. Because gatekeeping is presented as zero-sum—there is only one traffic entrance directly connecting to users—API gating and in-house agents cannot save superplatforms, leaving proactive adversarial attack as the most rational and inevitable choice. The paper explicitly labels this position as an envisioned trend to raise awareness, not as advocacy.
Load-bearing premise
The argument depends on the claim that only one traffic entrance gatekeeper can directly connect to users, so agent gatekeeping necessarily displaces platform gatekeeping; if users routinely interact through both layers at once, the necessity of attack collapses.
Editorial extensions
If this is right
- API gating ceases to be a sufficient defense: GUI agents that mimic human clicks and scrolling are largely immune to it, so superplatforms will invest in interface-level attacks.
- Adversarial research on agents will shift from white-box or task-specific attacks toward universal black-box attacks that obstruct any agent's task without knowing the user's instructions.
- If agents become the primary interface, superplatforms lose raw user behavioral data, which erodes their recommendation and advertising optimization even before revenue falls.
- The paper's forecast implies a race: whoever becomes the gatekeeper of user intent first—agent providers or platforms—will largely determine the future distribution of digital advertising and transaction revenue.
Reading between the lines
- The paper asserts, rather than proves, that only one traffic entrance gatekeeper can exist; a market model allowing users to split attention between direct platform use and agent-mediated use would test whether coexistence equilibria exist.
- One testable extension is to measure whether agent-mediated query volume and platform advertising revenue per user move inversely across several markets, which would support the zero-sum premise.
- Another implication the paper leaves implicit: if superplatforms do develop invisible attacks, regulation or technical standards for agent-platform interoperability would become as important as agent safety itself.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that LLM-driven AI agents pose a fundamental threat to superplatforms because they bypass user-attention-based monetization and can become the new gatekeepers of digital traffic. Applying gatekeeping theory, the authors conclude that superplatforms must proactively attack AI agents, especially GUI agents, to preserve their control. Section 2 develops the threat analysis and strategic reasoning; Section 3 presents a taxonomy of adversarial attacks on GUI agents organized by goals, attacker knowledge, attack visibility, and attack timing, together with challenges for superplatform-initiated attacks; Section 4 discusses and rebuts two alternative views (data moats and functional complementarity); Section 5 clarifies that the position is intended as a descriptive forecast rather than advocacy.
Significance. The paper identifies a strategically important and understudied intersection of AI agent security and platform economics. Its taxonomy of superplatform-initiated attacks in Section 3 is a useful organizational contribution, and it honestly engages with counterarguments rather than ignoring them. If the central thesis were established, the paper would point toward a new class of economically motivated, user-invisible attacks. However, the core modal claim — that superplatforms 'have to' attack agents — is not derived from a formal model, systematic evidence, or a rigorous decision-theoretic comparison. The contribution is therefore best characterized as a plausible position paper and research agenda, not a demonstrated result.
major comments (3)
- [§4.2] The claim 'since there is only one traffic entrance gatekeeper directly connecting to the users' is the load-bearing premise that converts a plausible threat into a necessity. This premise is asserted in a single sentence and is not derived from the cited gatekeeping literature (Lewin 1943; Shoemaker and Vos 2009; Barzilai-Nahon 2009), which describes gatekeeping processes but does not establish a single-gatekeeper uniqueness theorem. Actual digital markets contain layered intermediation — operating systems, app stores, search engines, and payment rails coexist and extract rents at different layers. If an agent can be the intent-layer gatekeeper while a superplatform retains the transaction/fulfillment/trust layer and both earn positive rents, the complementarity scenario discussed in §4.2 remains viable and the 'have to attack' conclusion does not follow. The paper should either provide a model or empirical evidence for the uniqueness premise, or weaken the conclusion to a conditional claim about incentives under a specified structural assumption.
- [§2.2] The inference from 'neither proprietary agents nor API gating is sufficient' to 'proactive adversarial attack is the most rational and inevitable choice' is a non sequitur. The comparison in §2.2 considers only three options: building proprietary agents, API gating, and proactive attack. It does not analyze other feasible responses such as revenue-sharing agreements with agents, agent-friendly APIs with margin extraction, differentiated premium services, improved human-facing experiences, regulatory lobbying, or the option of attacking only some agent classes. It also does not model the costs, success probabilities, or reputational risks of attacks, which are central to any 'most rational' conclusion. A decision-theoretic or game-theoretic comparison is needed to support the modal claim, or the claim should be reduced to 'one plausible strategy'.
- [§2.1] The empirical support for the claim that agents are displacing superplatforms as the traffic entrance is thin. The paper cites a single statement that Google's global search market share dropped below 90% 'for the first time since 2015' [1], sourced to a marketing trade article, plus two illustrative statistics on Google's advertising revenue and Amazon's daily transactions. This is insufficient to establish that AI agents are becoming the primary entrance or that the observed decline is causally attributable to agents rather than to other market changes. The paper should either add systematic data (e.g., time series of agent usage, platform traffic, and revenue across multiple platforms) or explicitly frame the central claim as a speculative forecast whose purpose is to stimulate discussion rather than to report an established empirical trend.
minor comments (5)
- [Abstract and front matter] There are several typos: 'perserving' should be 'preserving', 'Univeristy' should be 'University', 'Adersarial' in the Section 2.2 heading should be 'Adversarial', and reference [17] lists the author as 'V os' instead of 'Vos'.
- [References [1] and [18]] The paper relies on trade-press and aggregator sources for key quantitative claims; please cite the primary data sources or provide more rigorous references with clear access dates.
- [§3.2] The assertion that superplatforms' lack of knowledge is 'even more severe than typical black-box scenarios' is not justified; platforms can observe long histories of UI interactions, run A/B tests, and infer agent behavior from clickstream patterns, which is a form of query access that many black-box threat models do not include.
- [§3.5] The claim that superplatform-initiated attacks constitute a 'brand-new combination of adversarial attack attributes' is overstated; some cited environmental-injection attacks (e.g., [22], [14]) already combine black-box, user-invisible, and execution-phase properties in platform-like settings. The paper should clarify what is genuinely new: the business motivation, the universal obstruction goal, or the constraint of attacking through live platforms.
- [§5] The ethical disclaimer does not resolve the ambiguity between 'is' and 'ought'; even as a descriptive forecast, 'have to attack' can be read as a justification. The paper should more sharply separate the positive claim (platforms will be incentivized) from any normative implication, and this distinction should be maintained throughout the title and Section 2.
Circularity Check
No significant circularity: the paper's core claim is an argumentative forecast built on external gatekeeping theory and independent attack studies; its only self-citations are background references, and no prediction is fitted from its inputs.
full rationale
The derivation chain is: (1) superplatforms monetize attention by acting as traffic gatekeepers; (2) AI agents bypass that monetization and can become new gatekeepers; (3) therefore superplatforms have to attack agents. Each step is supported by external citations or asserted economic reasoning rather than by the paper's own outputs. The gatekeeping frame is imported from Lewin (1943), Shoemaker and Vos (2009), and Barzilai-Nahon (2009), and the adversarial-attack evidence is from independent empirical studies (e.g., Zhang et al. pop-ups, AdvWeb, BadAgent). There are no equations, fitted parameters, or benchmark predictions, so the fitted-input and self-definitional circularity patterns do not apply. The only self-citations are [24], a survey of AI agent protocols, and [29], on agentic information retrieval, used for background definitions of agent capabilities; those claims are not the load-bearing step that forces the attack conclusion. The main vulnerability of the paper is logical support, not circularity: Section 4.2 dismisses functional complementarity because 'there is only one traffic entrance gatekeeper directly connecting to the users,' a uniqueness premise that is asserted rather than derived. An unsupported premise is a correctness risk, not a circular reduction, because the conclusion is not equivalent to the input by definition; it simply rests on a contestable assumption. The Section 5 statement that the forecast is 'derived from our gatekeeping theory' is descriptive of the paper's own method and does not close a loop with any fitted result. Overall, the paper is self-contained as an opinion and analysis piece and does not reduce to its own inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption Gatekeeping theory, imported from communication studies, accurately describes superplatform business models and AI agent behavior.
- ad hoc to paper There is only one primary traffic entrance gatekeeper between users and services, making gatekeeping a zero-sum contest.
- domain assumption Superplatforms cannot reliably distinguish GUI agents from human users, so attacks must be user-invisible and occur at perception or execution phases.
Cite this review
Pith. "Pith review of Superplatforms Have to Attack AI Agents." pith.science (2026). https://pith.science/paper/YOPWJELL
@misc{pith2026250517861,
author = {Pith},
title = {Pith review of: Superplatforms Have to Attack AI Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/YOPWJELL}},
note = {Machine review of arXiv:2505.17861}
}
read the original abstract
Over the past decades, superplatforms, digital companies that integrate a vast range of third-party services and applications into a single, unified ecosystem, have built their fortunes on monopolizing user attention through targeted advertising and algorithmic content curation. Yet the emergence of AI agents driven by large language models (LLMs) threatens to upend this business model. Agents can not only free user attention with autonomy across diverse platforms and therefore bypass the user-attention-based monetization, but might also become the new entrance for digital traffic. Hence, we argue that superplatforms have to attack AI agents to defend their centralized control of digital traffic entrance. Specifically, we analyze the fundamental conflict between user-attention-based monetization and agent-driven autonomy through the lens of our gatekeeping theory. We show how AI agents can disintermediate superplatforms and potentially become the next dominant gatekeepers, thereby forming the urgent necessity for superplatforms to proactively constrain and attack AI agents. Moreover, we go through the potential technologies for superplatform-initiated attacks, covering a brand-new, unexplored technical area with unique challenges. We have to emphasize that, despite our position, this paper does not advocate for adversarial attacks by superplatforms on AI agents, but rather offers an envisioned trend to highlight the emerging tensions between superplatforms and AI agents. Our aim is to raise awareness and encourage critical discussion for collaborative solutions, prioritizing user interests and perserving the openness of digital ecosystems in the age of AI agents.
Figures
Reference graph
Works this paper leans on
-
[1]
From seo to geo: How agencies are navigat- ing llm-driven search
Campaign Asia. From seo to geo: How agencies are navigat- ing llm-driven search. https://www.campaignasia.com/article/ from-seo-to-geo-how-agencies-are-navigating-llm-driven-search/501593 ,
-
[2]
Gatekeeping: A critical review
Karine Barzilai-Nahon. Gatekeeping: A critical review. Annual Review of Information Science and Technology, 43(1):1–79, 2009
work page 2009
-
[3]
The obvious invisible threat: Llm-powered gui agents’ vulnerability to fine-print injections
Chaoran Chen, Zhiping Zhang, Bingcan Guo, Shang Ma, Ibrahim Khalilov, Simret A Gebreegzi- abher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, et al. The obvious invisible threat: Llm-powered gui agents’ vulnerability to fine-print injections. arXiv preprint arXiv:2504.11281, 2025
arXiv 2025
-
[5]
Yurun Chen, Xueyu Hu, Keting Yin, Juncheng Li, and Shengyu Zhang. Aeia-mn: Evaluating the robustness of multimodal llm-powered mobile agents against active environmental injection attacks. arXiv preprint arXiv:2502.13053, 2025
arXiv 2025
-
[6]
Exploring the adversarial robustness of clip for ai-generated image detection
Vincenzo De Rosa, Fabrizio Guillaro, Giovanni Poggi, Davide Cozzolino, and Luisa Verdoliva. Exploring the adversarial robustness of clip for ai-generated image detection. In 2024 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–6. IEEE, 2024
work page 2024
-
[7]
Ali Dorri, Salil S Kanhere, and Raja Jurdak. Multi-agent systems: A survey. Ieee Access, 6: 28573–28593, 2018
work page 2018
- [8]
-
[9]
Clip-guided generative networks for transferable targeted adversarial attacks
Hao Fang, Jiawei Kong, Bin Chen, Tao Dai, Hao Wu, and Shu-Tao Xia. Clip-guided generative networks for transferable targeted adversarial attacks. In European Conference on Computer Vision, pages 1–19. Springer, 2024
work page 2024
Show all 32 references
-
[10]
Resist platform-controlled ai agents and champion user-centric agent advocates
Sayash Kapoor, Noam Kolt, and Seth Lazar. Resist platform-controlled ai agents and champion user-centric agent advocates. arXiv preprint arXiv:2505.04345, 2025
2025 arXiv
-
[11]
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju. Certifying llm safety against adversarial prompting. arXiv preprint arXiv:2309.02705, 2023
2023 arXiv
-
[12]
Amazon statistics you should know in 2024
LandingCube. Amazon statistics you should know in 2024. https://landingcube.com/ amazon-statistics/?utm_source=chatgpt.com, 2024. Accessed: 2025-05-23
2024
-
[13]
Forces behind food habits and methods of change
Kurt Lewin. Forces behind food habits and methods of change. Bulletin of the National Research Council, (108):35–65, 1943
1943
-
[14]
Eia: Environmental injection attack on generalist web agents for privacy leakage
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun. Eia: Environmental injection attack on generalist web agents for privacy leakage. arXiv preprint arXiv:2409.11295, 2024
2024 arXiv
-
[15]
Caution for the environment: Multimodal agents are susceptible to environmental distractions
Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, and Hai Zhao. Caution for the environment: Multimodal agents are susceptible to environmental distractions. arXiv preprint arXiv:2408.02544, 2024. 10
2024 arXiv
-
[16]
Future internet and digital ecosystems
Tiziana Russo Spena, Marco Tregua, and Francesco Bifulco. Future internet and digital ecosystems. Digital Transformation in the Cultural Heritage Sector: Challenges to Marketing in the New Digital Era, pages 17–38, 2021
2021
-
[17]
Shoemaker and Timothy P
Pamela J. Shoemaker and Timothy P. V os.Gatekeeping Theory. Routledge, New York, 2009. ISBN 9780415981392
2009
-
[18]
Advertising revenue of google from 2001 to 2023
Statista. Advertising revenue of google from 2001 to 2023. https://www.statista.com/ statistics/266249/advertising-revenue-of-google/ , 2024. Accessed: 2025-05-23
2001
-
[19]
Badagent: Inserting and activating backdoor attacks in llm agents
Yifei Wang, Dizhan Xue, Shengjie Zhang, and Shengsheng Qian. Badagent: Inserting and activating backdoor attacks in llm agents. arXiv preprint arXiv:2406.03007, 2024
2024 arXiv
-
[20]
Foot-in-the-door: A multi-turn jailbreak for llms, 2025
Zixuan Weng, Xiaolong Jin, Jinyuan Jia, and Xiangyu Zhang. Foot-in-the-door: A multi-turn jailbreak for llms, 2025. URL https://arxiv.org/abs/2502.19820
2025 arXiv
-
[21]
Dissecting adversarial robustness of multimodal lm agents
Chen Henry Wu, Rishi Rajesh Shah, Jing Yu Koh, Russ Salakhutdinov, Daniel Fried, and Aditi Raghunathan. Dissecting adversarial robustness of multimodal lm agents. In The Thirteenth International Conference on Learning Representations
-
[22]
Advweb: Controllable black-box attacks on vlm-powered web agents, 2024
Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li. Advweb: Controllable black-box attacks on vlm-powered web agents, 2024. URL https://arxiv.org/abs/2410.17401
2024 arXiv
-
[23]
Watch out for your agents! investigating backdoor threats to llm-based agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. Watch out for your agents! investigating backdoor threats to llm-based agents. Advances in Neural Information Processing Systems, 37:100938–100964, 2024
2024
-
[24]
A survey of ai agent protocols
Yingxuan Yang, Huacan Chai, Yuanyi Song, Siyuan Qi, Muning Wen, Ning Li, Junwei Liao, Haoyi Hu, Jianghao Lin, Gaowei Chang, et al. A survey of ai agent protocols. arXiv preprint arXiv:2504.16736, 2025
2025 arXiv
-
[25]
The dawn of lmms: Preliminary explorations with gpt-4v(ision), 2023
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The dawn of lmms: Preliminary explorations with gpt-4v(ision), 2023. URL https://arxiv.org/abs/2309.17421
2023 arXiv
-
[26]
Infecting llm agents via generalizable adversarial attack
Weichen Yu, Kai Hu, Tianyu Pang, Chao Du, Min Lin, and Matt Fredrikson. Infecting llm agents via generalizable adversarial attack. In Red Teaming GenAI: What Can We Learn from Adversaries?
-
[27]
Infecting LLM agents via generalizable adversarial attack
Weichen Yu, Kai Hu, Tianyu Pang, Chao Du, Min Lin, and Matt Fredrikson. Infecting LLM agents via generalizable adversarial attack. In Red Teaming GenAI: What Can We Learn from Adversaries?, 2025. URL https://openreview.net/forum?id=udsmFGMwlp
2025
-
[28]
Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents, 2025
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents, 2025. URL https://arxiv.org/abs/2410.02644
2025 arXiv
-
[29]
Agentic information retrieval
Weinan Zhang, Junwei Liao, Ning Li, Kounianhua Du, and Jianghao Lin. Agentic information retrieval. arXiv preprint arXiv:2410.09713, 2024
2024 arXiv
-
[30]
Attacking vision-language computer agents via pop-ups
Yanzhe Zhang, Tao Yu, and Diyi Yang. Attacking vision-language computer agents via pop-ups. arXiv preprint arXiv:2411.02391, 2024
2024 arXiv
-
[31]
Qava: Query-agnostic visual attack to large vision-language models
Yudong Zhang, Ruobing Xie, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, and Yu Wang. Qava: Query-agnostic visual attack to large vision-language models. arXiv preprint arXiv:2504.11038, 2025
2025 arXiv
-
[32]
Data poisoning attacks on multi-task relationship learning
Mengchen Zhao, Bo An, Yaodong Yu, Sulin Liu, and Sinno Pan. Data poisoning attacks on multi-task relationship learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018. 11
2018
-
[2024]
Accessed: 2025-05-23
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.