REVIEW 2 major objections 3 minor 1 cited by
A Red Teaming Roadmap Towards System-Level Safety
T0 review · 2 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read LLM red teaming should target product safety specs and system-level defenses, not abstract harms.
desk verdict A clear, well-organized position paper that reprioritizes red-teaming research; its central empirical assumption about where real-world harms live is undefended, but the roadmap is still worth engaging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the three-level distinction among model, product, and system, together with the product safety specification as the red teaming objective. A product safety specification states what behavior is prohibited for that specific deployment, including who the users are, what tools are attached, and what regulations apply, and the paper treats it as the target that a successful red team finding must violate. The second mechanism is the system-level mitigation loop: trajectory monitoring classifies harm over whole interaction histories, user monitoring detects and bans repeat malicious users, and rapid response patches newly found failures before they are exploited at scale. Red teaming the monitor, through sabotage experiments where an adversary must complete a harmful task without detection, stress-tests that loop. These mechanisms convert the vague question of whether a model is safe into testable questions about a specific product and its deployed defenses.
What would settle it
A head-to-head deployment study would settle it: take the same underlying model in the same product, run one arm with only refusal training and one arm whose red teaming targets the product safety specification with system-level monitoring and user banning, and compare measured harm per active user, for example substantiated abuse reports per 100,000 interactions. If the model-only arm shows equal or lower harm, the paper's priority claim is false.
Extended reading notes
Core claim
The paper's central claim is that LLM red teaming will reduce real-world harm only if it is re-anchored around two priorities: product safety specifications over abstract social biases or ethical principles, and system-level safety over model-level robustness. It defines a model as the neural network, a product as the deployed application built on it, and a system as the product plus its deployment infrastructure, including monitors, staff, users, and environment, and argues that the attack surface users actually face is the end-to-end product stack, not the bare model. Because safety is contextual, a behavior that violates one product's policy may be acceptable in another, so red teaming objectives should be explicit, actionable, and measured against the target product's stated spec. The paper then argues that realistic threat models differ by product type, with multi-turn conversations for chatbots, prosodic and multilingual channels for audio assistants, frame-spanning harm for video generators, and tool- and environment-driven attacks for agents, and that safeguards should be tested both independently and in realistic sandboxes. Its final step is the claim that system-level measures such as trajectory monitoring, user monitoring, rapid response to newly discovered jailbreaks, and red teaming the monitor itself through sabotage experiments are necessary to make red teaming relevant to deployed risk.
Load-bearing premise
The roadmap assumes that a product safety specification can be written precisely enough to serve as a red teaming objective, and that red teaming against that spec plus system-level mitigation reduces real-world harm more effectively than model-level testing.
Editorial extensions
If this is right
- Red teaming benchmarks would be evaluated against a stated product safety specification and deployment context, so a jailbreak that violates no deployed product's policy would count as low priority.
- Multi-turn and trajectory-level attacks would receive more research attention than single-turn jailbreaks, because defenses trained for single turns do not generalize to extended conversations.
- Safety investments would shift toward asynchronous monitoring, malicious-user detection, and rapid safeguard deployment, since most harm is diffuse over many requests rather than a single catastrophic output.
- White-box red teaming would assume the adversary can fine-tune the model, and black-box red teaming would assume limited query access, making attacks match real attacker resources.
- Agent red teaming would require investment in realistic sandboxes and simulated environments, because tool outputs and environmental context are part of the attack surface.
Reading between the lines
- Editorial inference: If product-specific specifications become the primary evaluation target, standardized public benchmarks may fragment, because each product's policy is different; a shared harm taxonomy would survive only where regulators impose one.
- Editorial inference: The emphasis on system-level monitoring implies that model-level alignment research could be deprioritized, yet the system-level loop still depends on the model not producing a single catastrophic output before a monitor can intervene, so the two are complements rather than substitutes.
- Editorial inference: Red teaming the monitor sets up an arms race between detection evasion and detection improvement; sabotage experiments would need to be repeated continuously as monitors update, making red teaming an ongoing operational process rather than a one-time evaluation.
- Editorial inference: A testable consequence is that products with user banning and rapid patching should show lower measured abuse per user than products with only refusal-trained models, holding the underlying model fixed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that LLM red teaming research currently misallocates effort. It recommends three priorities: (1) red teaming should target product safety specifications rather than abstract social biases or ethical principles; (2) threat models should reflect realistic attackers and deployment contexts rather than idealized single-query settings; and (3) red teaming should expand from model-level refusal robustness to system-level safety, including trajectory monitoring, user monitoring, rapid patching, and red teaming of monitors. The paper motivates these claims through a taxonomy of model/product/system, worked threat-model sketches for chatbots, audio assistants, video generators, and agents, and a set of best practices in Section 4. Section 5 develops the system-level argument, and Section 6 considers alternative views. The paper is explicitly a roadmap rather than an empirical study; its central claims are normative and strategic.
Significance. If the roadmap is accepted, it would redirect a substantial portion of red teaming research from universal jailbreak benchmarks toward deployment-aware evaluation and system-level defenses. The paper's contribution is mainly organizational and rhetorical: it clarifies distinctions among model, product, and system; identifies understudied attack surfaces such as multi-turn conversations, audio, video, agents, and tool outputs; and offers a practical best-practices list in Section 4. It is honest about its status as a position paper, explicitly engaging with an alternative view in Section 6. The main weakness is that the two prioritizations—product over abstract, and system over model—are supported by selected examples rather than by systematic evidence, leaving the strongest comparative claims underdetermined. As a position paper, it need not contain new experiments, but its empirical premises should be supported with a structured literature synthesis or measurement study.
major comments (2)
- [Section 5, 'Trajectory and User Monitoring' and 'Rapid Response'] The central comparative claim that system-level safety should be prioritized over model-level robustness rests on the empirical premise that 'many realistic harmful uses of an AI system may require its cooperation over multiple turns' and that asynchronous monitoring can catch failures before harm scales. The manuscript provides examples and citations (e.g., [30, 47, 56]) but no evidence about the relative frequency or severity of diffuse multi-turn harms versus single-output catastrophic harms. The introduction's motivating examples—dangerous biological agents and other catastrophic risks—are precisely the kind of single-response harms for which post-hoc user banning and patching cannot undo the first harmful output. The sentence 'as long as sufficiently effective asynchronous monitoring detects these failures as they happen' is a conditional, not a demonstrated premise, and it is hardest for the highest-severity cases. To support the prioritization, the paper should either supply or cite data on harm distributions and intervention lead times in deployed products, or soften the claim from 'should be prioritized over' to 'should complement.'
- [Section 2, 'Safety Specifications Should Focus on Product-Level Risk'] The first pillar assumes that a product safety specification can be defined precisely enough to serve as a red teaming objective and that such specifications are tied to actual societal harm. The paper acknowledges that products can and should have divergent safety considerations and that many products lack well-defined specifications; its fallback is that researchers should 'define or infer a plausible policy context.' This fallback reintroduces the same normative ambiguity that the paper criticizes in abstract harm categories, because a plausible policy may be vague or contested. The manuscript does not present a worked example in which a product specification is translated into a concrete red teaming metric (what counts as a successful breach, what the evaluation metric is) for a specific product. Without at least one such case study, the claim that product-level red teaming is more tractable and more decision-relevant than model-level harm testing remains an assertion. A worked example would make the proposal easier to adopt and would make its central premise falsifiable.
minor comments (3)
- [Section 4, 'Assessing the Delta'] The sentence 'There would be no difference in the models vulnerabilities' contains a typo; it should read 'the model's vulnerabilities.'
- [Section 3.1, first bullet list] The claim that 'more difficult single-turn attacks seem to be independent, ad-hoc issues with safeguards that can be quickly patched' is supported by a single self-cited reference [63]; a systematic comparison or an independent replication would make the argument more robust.
- [Section 6, 'Argument for Static Harm Research'] The alternative view is presented fairly but dismissed in a few sentences; readers who disagree with the product-specification premise would benefit from a more detailed rebuttal, particularly concerning the role of regulation and societal consensus.
Circularity Check
No circularity: the paper is a position piece whose recommendations rest on contextual arguments and conditional empirical claims, not on self-citations or fitted quantities.
full rationale
This is a position paper with no equations, fitted parameters, or formal derivation chain. Its central recommendation—prioritize product safety specifications and system-level safeguards—is argued from definitions and stated value judgments, not reduced to its own inputs. The paper does cite several works from its own team (e.g., [47] on multi-turn jailbreaks, [63] on rapid response, and [41]/[90] on agent and browser red teaming), but these are supporting examples for empirical premises that are also supported by independent citations or are externally testable; the recommendations do not reduce to those citations. The load-bearing conditional about asynchronous monitoring is explicitly conditional and is not presented as a demonstrated premise. No self-definitional step, fitted input, imported uniqueness theorem, or ansatz-smuggling chain is present. Concerns about unquantified harm distributions are evidence gaps or correctness risks, not circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption Product safety specifications can be defined precisely enough to serve as measurable red teaming objectives.
- domain assumption Threat models that approximate real attackers are more useful for mitigating real-world harm than abstract or idealized models.
- domain assumption System-level mitigations such as monitoring, user banning, and rapid patching are available and effective for deployed AI products.
- domain assumption Single-turn jailbreak failures are largely ad-hoc and patchable, while multi-turn and system-level harms are more fundamental.
- domain assumption Current conference red teaming submissions do not prioritize the right research problems.
Cite this review
Pith. "Pith review of A Red Teaming Roadmap Towards System-Level Safety." pith.science (2026). https://pith.science/paper/Q6C3I4FI
@misc{pith2026250605376,
author = {Pith},
title = {Pith review of: A Red Teaming Roadmap Towards System-Level Safety},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q6C3I4FI}},
note = {Machine review of arXiv:2506.05376}
}
read the original abstract
Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has effectively identified critical vulnerabilities in state-of-the-art refusal-trained LLMs. However, in our view the many conference submissions on LLM red teaming do not, in aggregate, prioritize the right research problems. First, testing against clear product safety specifications should take a higher priority than abstract social biases or ethical principles. Second, red teaming should prioritize realistic threat models that represent the expanding risk landscape and what real attackers might do. Finally, we contend that system-level safety is a necessary step to move red teaming research forward, as AI models present new threats as well as affordances for threat mitigation (e.g., detection and banning of malicious users) once placed in a deployment context. Adopting these priorities will be necessary in order for red teaming research to adequately address the slate of new threats that rapid AI advances present today and will present in the very near future.
Figures
Forward citations
Cited by 1 Pith paper
-
Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
Across 2M+ in-the-wild LLM conversations, jailbreak attempts show no higher complexity than normal chats, and assistant toxicity has declined over time, suggesting bounded attack sophistication.
Reference graph
Works this paper leans on
-
[1]
URL https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995
Anthropic. URL https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995. pdf
-
[2]
Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku
Anthropic_Computer_Use. Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku. URLhttps://www.anthropic.com/news/3-5-models-and-computer-use
-
[3]
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield- Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan. ...
arXiv 2022
-
[4]
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran- Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. ...
arXiv 2022
- [5]
-
[6]
B. R. Bartoldson, J. Diffenderfer, K. Parasyris, and B. Kailkhura. Adversarial robustness limits via scaling-law and human-alignment studies.arXiv preprint arXiv:2404.09349, 2024
arXiv 2024
-
[7]
Y. Bengio, G. Hinton, A. Yao, D. Song, P . Abbeel, S. Russell, P . Torr, J. Brauner, S. Mindermann, et al. Managing extreme AI risks amid rapid progress.Science, 384(6698):916–919, 2024. doi: 10.1126/science.adn0117
-
[8]
J. Benton, M. Wagner, E. Christiansen, C. Anil, E. Perez, J. Srivastav, E. Durmus, D. Ganguli, S. Kravec, B. Shlegeris, J. Kaplan, H. Karnofsky, E. Hubinger, R. Grosse, S. R. Bowman, and D. Duvenaud. Sabotage evaluations for frontier models, 2024. URL https://arxiv.org/abs/ 2410.21514
arXiv 2024
Show all 111 references
-
[9]
A. N. Bhagoji, W. He, B. Li, and D. Song. Practical black-box attacks on deep neural networks using efficient query mechanisms. InProceedings of the European conference on computer vision (ECCV), pages 154–169, 2018
2018
-
[10]
Bonatti, D
R. Bonatti, D. Zhao, F. Bonacci, D. Dupont, S. Abdali, Y. Li, Y. Lu, J. Wagle, K. Koishida, A. F. C. Bucker, L. Jang, and Z. Hui. Windows agent arena: Evaluating multi-modal os agents at scale. ArXiv, abs/2409.08264, 2024. URLhttps://api.semanticscholar.org/CorpusID:272600411. 11
2024 arXiv
-
[11]
Carlini and D
N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. In2017 ieee symposium on security and privacy (sp), pages 39–57. Ieee, 2017
2017
-
[12]
P . Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong. Jailbreaking black box large language models in twenty queries.arXiv preprint arXiv:2310.08419, 2023. URL https: //arxiv.org/abs/2310.08419
2023 arXiv
-
[13]
C. Chen, Z. Zhang, B. Guo, S. Ma, I. Khalilov, S. A. Gebreegziabher, Y. Ye, Z. Xiao, Y. Yao, T. Li, and T. J.-J. Li. The obvious invisible threat: Llm-powered gui agents’ vulnerability to fine-print injections, 2025. URLhttps://arxiv.org/abs/2504.11281
2025 arXiv
-
[14]
G. Chen, F. Song, Z. Zhao, X. Jia, Y. Liu, Y. Qiao, and W. Zhang. Audiojailbreak: Jailbreak attacks against end-to-end large audio-language models, 2025. URL https://arxiv.org/abs/ 2505.14103
2025
-
[15]
P .-Y. Chen, H. Zhang, Y. Sharma, J. Yi, and C.-J. Hsieh. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. InProceedings of the 10th ACM workshop on artificial intelligence and security, pages 15–26, 2017
2017
-
[16]
Cheng, C
G. Cheng, C. Zhang, W. Cai, L. Zhao, C. Sun, and J. Bian. Empowering large language models on robotic manipulation with affordance prompting, 2024. URL https://arxiv.org/abs/2404. 11027
2024
-
[17]
Croce and M
F. Croce and M. Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. InICML, 2020
2020
-
[18]
Croce, M
F. Croce, M. Andriushchenko, V . Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P . Mittal, and M. Hein. Robustbench: a standardized adversarial robustness benchmark.arXiv preprint arXiv:2010.09670, 2020
2010 arXiv
-
[19]
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu. MASTERKEY: Automated jailbreaking of large language model chatbots. InProceedings of the 2024 Network and Distributed System Security Symposium (NDSS), 2024
2024
-
[20]
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su. Mind2web: Towards a generalist agent for the web.ArXiv, abs/2306.06070, 2023. URL https://api.semanticscholar. org/CorpusID:259129428
2023 arXiv
-
[21]
For-profit AI safety: AI safety needs to scale and here’s how you can do it
Esben Kran. For-profit AI safety: AI safety needs to scale and here’s how you can do it. https:// apartresearch.com/news/ai-safety-needs-to-scale-and-heres-how-you-can-do-it , 2024. Apart Research blog, accessed 22 May 2025
2024
-
[22]
P . M. Gade, S. Lermen, C. Rogers-Smith, and J. Ladish. Badllama: cheaply removing safety fine- tuning from llama 2-chat 13b.ArXiv, abs/2311.00117, 2023. URL https://api.semanticscholar. org/CorpusID:264832925
2023 arXiv
-
[23]
Ghosh, P
S. Ghosh, P . Varshney, E. Galinkin, and C. Parisien. Aegis: Online adaptive ai content safety moderation with ensemble of llm experts.ArXiv, abs/2404.05993, 2024. URL https://api. semanticscholar.org/CorpusID:269009460
2024 arXiv
-
[24]
I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples.arXiv preprint arXiv:1412.6572, 2014
2014 arXiv
-
[25]
Google. Veo. URLhttps://deepmind.google/models/veo/. 12
-
[26]
Greenblatt, B
R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger. Ai control: Improving safety despite intentional subversion, 2024. URLhttps://arxiv.org/abs/2312.06942
2024 arXiv
-
[27]
M. Y. Guan, M. Joglekar, E. Wallace, S. Jain, B. Barak, A. Helyar, R. Dias, A. Vallone, H. Ren, J. Wei, H. W. Chung, S. Toyer, J. Heidecke, A. Beutel, and A. Glaese. Deliberative alignment: Reasoning enables safer language models, 2025. URLhttps://arxiv.org/abs/2412.16339
2025 arXiv
-
[28]
Z. Guan, M. Hu, R. Zhu, S. Li, and A. Vullikanti. Benign samples matter! fine-tuning on outlier benign samples severely breaks safety. 2025. URL https://api.semanticscholar.org/CorpusID: 278501392
2025
-
[29]
Hafez, A
A. Hafez, A. N. Akhormeh, A. Hegazy, and A. Alanwar. Safe llm-controlled robots with formal guarantees via reachability analysis, 2025. URLhttps://arxiv.org/abs/2503.03911
2025 arXiv
-
[30]
Automated multi-turn red-teaming with cascade
Haize. Automated multi-turn red-teaming with cascade. URL https://www.haizelabs.com/ technology/automated-multi-turn-red-teaming-with-cascade
-
[31]
Hammond, A
L. Hammond, A. Chan, J. Clifton, J. Hoelscher-Obermaier, A. Khan, E. McLean, C. Smith, W. Bar- fuss, J. Foerster, T. Gavenˇ ciak, T. A. Han, E. Hughes, V . Kovaˇ rík, J. Kulveit, J. Z. Leibo, C. Oester- held, C. S. de Witt, N. Shah, M. Wellman, P . Bova, T. Cimpeanu, C. Ezell,...
2025
-
[32]
W. Held, M. J. Ryan, A. Shrivastava, A. S. Khan, C. Ziems, E. Li, M. Bartelds, M. Sun, T. Li, W. Gan, and D. Yang. Cava: Comprehensive assessment of voice assistants. https://github. com/SALT-NLP/CAVA, 2025. URL https://talkarena.org/cava. A benchmark for evaluating large audi...
2025
-
[33]
Hendrycks, M
D. Hendrycks, M. Mazeika, and T. Woodside. An overview of catastrophic ai risks, 2023. URL https://arxiv.org/abs/2306.12001
2023 arXiv
-
[34]
Huang, S
T. Huang, S. Hu, F. Ilhan, S. F. Tekin, and L. Liu. Harmful fine-tuning attacks and defenses for large language models: A survey, 2024. URLhttps://arxiv.org/abs/2409.18169
2024 arXiv
-
[35]
Hughes, S
J. Hughes, S. Price, A. Lynch, R. Schaeffer, F. Barez, S. Koyejo, H. Sleight, E. Jones, E. Perez, and M. Sharma. Best-of-n jailbreaking, 2024. URLhttps://arxiv.org/abs/2412.03556
2024 arXiv
-
[36]
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa. Llama guard: Llm-based input-output safeguard for human-ai conversations,
-
[37]
Jones, A
E. Jones, A. Dragan, and J. Steinhardt. Adversaries can misuse combinations of safe models, 2024. URLhttps://arxiv.org/abs/2406.14595
2024 arXiv
-
[38]
J. Y. Koh, R. Lo, L. Jang, V . Duvvur, M. C. Lim, P .-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.ArXiv, abs/2401.13649, 2024. URLhttps://api.semanticscholar.org/CorpusID:267199749
2024 arXiv
-
[39]
Kritz, V
J. Kritz, V . Robinson, R. Vacareanu, B. Varjavand, M. Choi, B. Gogov, S. R. Team, S. Yue, W. E. Primack, and Z. Wang. Jailbreaking to jailbreak, 2025. URL https://arxiv.org/abs/2502.09638
2025 arXiv
-
[40]
Krizhevsky, V
A. Krizhevsky, V . Nair, and G. Hinton. Cifar-10 (canadian institute for advanced research). URL http://www.cs.toronto.edu/~kriz/cifar.html. 13
-
[41]
Kumar, E
P . Kumar, E. Lau, S. Vijayakumar, T. Trinh, S. R. Team, E. Chang, V . Robinson, S. Hendryx, S. Zhou, M. Fredrikson, S. Yue, and Z. Wang. Refusal-trained llms are easily jailbroken as browser agents,
-
[42]
M. Kuo, J. Zhang, A. Ding, Q. Wang, L. DiValentin, Y. Bao, W. Wei, H. Li, and Y. Chen. H- cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash thinking, 2025. URL https: //...
2025 arXiv
-
[43]
Laban, H
P . Laban, H. Hayashi, Y. Zhou, and J. Neville. Llms get lost in multi-turn conversation, 2025. URL https://arxiv.org/abs/2505.06120
2025 arXiv
-
[44]
H. Lai, X. Liu, I. L. Iong, S. Yao, Y. Chen, P . Shen, H. Yu, H. Zhang, X. Zhang, Y. Dong, and J. Tang. Autowebglm: A large language model-based web navigating agent.Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024. URL https: //api.se...
2024
-
[45]
Le Jeune, J
P . Le Jeune, J. Liu, L. Rossi, and M. Dora. Realharm: A collection of real-world language model application failures.arXiv preprint arXiv:2504.10277, 2025
2025 arXiv
-
[46]
A. Li, Y. Zhou, V . C. Raghuram, T. Goldstein, and M. Goldblum. Commercial llm agents are already vulnerable to simple yet dangerous attacks, 2025. URLhttps://arxiv.org/abs/2502.08586
2025 arXiv
-
[47]
N. Li, Z. Han, I. Steneker, W. Primack, R. Goodside, H. Zhang, Z. Wang, C. Menghini, and S. Yue. Llm defenses are not robust to multi-turn human jailbreaks yet, 2024. URL https: //arxiv.org/abs/2408.15221
2024 arXiv
-
[48]
J. Liu, S. Liang, S. Zhao, R. Tu, W. Zhou, X. Cao, D. Tao, and S. K. Lam. Jailbreaking the text-to-video generative models, 2025. URLhttps://arxiv.org/abs/2505.06679
2025 arXiv
-
[49]
X. Liu, Z. Yu, Y. Zhang, N. Zhang, and C. Xiao. Automatic and universal prompt injection attacks against large language models.arXiv preprint arXiv:2403.04957, 2024
2024 arXiv
-
[50]
Y. Lu, T. Ju, M. Zhao, X. Ma, Y. Guo, and Z. Zhang. Eva: Red-teaming gui agents via evolving indirect prompt injection, 2025. URLhttps://arxiv.org/abs/2505.14289
2025 arXiv
-
[51]
Madry, A
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks.arXiv preprint arXiv:1706.06083, 2017
2017 arXiv
-
[52]
Mazeika, L
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
2024
-
[53]
A. MCP . Introducing the model context protocol. URL https://www.anthropic.com/news/ model-context-protocol
-
[54]
Mehrotra, M
A. Mehrotra, M. Zampetakis, P . Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi. Tree of attacks: Jailbreaking black-box llms automatically.arXiv preprint arXiv:2312.02119, 2023. URL https://arxiv.org/abs/2312.02119
2023 arXiv
-
[55]
Y. Miao, Y. Zhu, Y. Dong, L. Yu, J. Zhu, and X.-S. Gao. T2vsafetybench: Evaluating the safety of text-to-video generative models, 2024. URLhttps://arxiv.org/abs/2407.05965
2024 arXiv
-
[56]
Nakash, G
I. Nakash, G. Kour, G. Uziel, and A. Anaby-Tavor. Breaking react agents: Foot-in-the-door attack will get you in, 2024. URLhttps://arxiv.org/abs/2410.16950. 14
2024 arXiv
-
[57]
Narodytska and S
N. Narodytska and S. P . Kasiviswanathan. Simple black-box adversarial perturbations for deep networks.arXiv preprint arXiv:1612.06299, 2016
2016 arXiv
-
[58]
Nguyen, S
V .-A. Nguyen, S. Zhao, G. Dao, R. Hu, Y. Xie, and L. A. Tuan. Three minds, one legend: Jailbreak large reasoning model with adaptive stacked ciphers, 2025. URL https://arxiv.org/abs/2505. 16241
2025
-
[59]
Sora: Creating video from text
OpenAI. Sora: Creating video from text. https://openai.com/index/sora/, 2024. Ac- cessed 22 May 2025
2024
-
[60]
URLhttps://cdn.openai.com/operator_system_card.pdf
Operator. URLhttps://cdn.openai.com/operator_system_card.pdf
-
[61]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P . Welinder, P . Christiano, J. Leike, and R. Lowe. Training language models to follow instructio...
2022 arXiv
-
[62]
Papernot, P
N. Papernot, P . McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami. Practical black-box attacks against machine learning. InProceedings of the 2017 ACM on Asia conference on computer and communications security, pages 506–519, 2017
2017
-
[63]
A. Peng, J. Michael, H. Sleight, E. Perez, and M. Sharma. Rapid response: Mitigating llm jailbreaks with a few examples, 2024. URLhttps://arxiv.org/abs/2411.07494
2024 arXiv
-
[64]
Perez, S
E. Perez, S. Ringer, K. Lukoši¯ut˙e, K. Nguyen, C. Pettit, C. Olsson, and et al. Discovering language model behaviors with model-written evaluations. InFindings of ACL 2023, pages 3419–3448, 2023
2023
-
[65]
S. Pichai. Cloud next 2023: Sharing the best of our AI with google cloud. https://blog.google/ products/google-cloud/cloud-next-2023-sundar-pichai-keynote/ , 2023. Google Blog, ac- cessed 22 May 2025
2023
-
[66]
X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P . Mittal, and P . Henderson. Safety alignment should be made more than just a few tokens deep, 2024. URL https://arxiv.org/abs/2406. 05946
2024
-
[67]
X. Qi, Y. Zeng, T. Xie, P .-Y. Chen, R. Jia, P . Mittal, and P . Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=hTEGyKf0dZ
2024
-
[68]
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y.-T. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, R. Tian, R. Xie, J. Zhou, M. H. Gerstein, D. Li, Z. Liu, and M. Sun. Toolllm: Facilitating large language models to master 16000+ real-world apis.ArXiv, abs/2307.16789, 2023. URL htt...
2023 arXiv
-
[69]
Radosevich and J
B. Radosevich and J. Halloran. Mcp safety audit: Llms with the model context protocol allow major security exploits, 2025. URLhttps://arxiv.org/abs/2504.03767
2025 arXiv
-
[70]
Rando, J
J. Rando, J. Zhang, N. Carlini, and F. Tramèr. Adversarial ml problems are getting harder to solve and to evaluate, 2025. URLhttps://arxiv.org/abs/2502.02260
2025 arXiv
-
[71]
M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, R. Comanescu, C. Akbulut, T. Stepleton, J. Mateos-Garcia, S. Bergman, J. Kay, C. Griffin, B. Bariach, I. Gabriel, V . Rieser, W. Isaac, and L. Weidinger. Gaps in the safety evaluation of generative ai. InProceedings of the 2024...
2024
-
[72]
Q. Ren, H. Li, D. Liu, Z. Xie, X. Lu, Y. Qiao, L. Sha, J. Yan, L. Ma, and J. Shao. Derail yourself: Multi-turn llm jailbreak attack through self-discovered clues, 2024. URL https://arxiv.org/abs/ 2410.10700
2024
-
[73]
Robey, Z
A. Robey, Z. Ravichandran, V . Kumar, H. Hassani, and G. J. Pappas. Jailbreaking llm-controlled robots.arXiv preprint arXiv:2410.13691, 2024
2024 arXiv
-
[74]
J. Roh, V . Shejwalkar, and A. Houmansadr. Multilingual and multi-accent jailbreaking of audio llms, 2025. URLhttps://arxiv.org/abs/2504.01094
2025 arXiv
-
[75]
Russinovich, A
M. Russinovich, A. Salem, and R. Eldan. Great, now write an article about that: The crescendo multi-turn llm jailbreak attack.arXiv preprint arXiv:2404.01833, 2024. URL https://arxiv.org/ abs/2404.01833
2024 arXiv
-
[76]
Sharma, M
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, E. Perez, and et al. Towards understanding sycophancy in language models. InProceedings of ICLR 2024, 2024
2024
-
[77]
Sharma, M
M. Sharma, M. Tong, J. Mu, J. Wei, J. Kruthoff, S. Goodfriend, E. Ong, A. Peng, R. Agarwal, C. Anil, A. Askell, N. Bailey, J. Benton, E. Bluemke, S. R. Bowman, E. Christiansen, H. Cunningham, A. Dau, A. Gopal, R. Gilson, L. Graham, L. Howard, N. Kalra, T. Lee, K. Lin, P . Lofg...
2025 arXiv
-
[78]
Sheshadri, A
A. Sheshadri, A. Ewart, P . Guo, A. Lynch, C. Wu, V . Hebbar, H. Sleight, A. C. Stickland, E. Perez, D. Hadfield-Menell, and S. Casper. Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms.arXiv preprint arXiv:2407.15549, 2024
2024 arXiv
-
[79]
Sikorski, L
P . Sikorski, L. Schrader, K. Yu, L. Billadeau, J. Meenakshi, N. Mutharasan, F. Esposito, H. AliAkbar- pour, and M. Babaiasl. Deployment of large language models to control mobile robots at the edge,
-
[80]
Y. Song, F. Xu, S. Zhou, and G. Neubig. Beyond browsing: Api-based web agents, 2025. URL https://arxiv.org/abs/2410.16464
2025 arXiv
-
[81]
Souly, Q
A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P . Abbeel, J. Svegliato, S. Emmons, O. Watkins, and S. Toyer. A strongreject for empty jailbreaks, 2024
2024
-
[82]
URLhttps://arxiv.org/abs/2405.17670
-
[83]
G. Sun, X. Zhan, S. Feng, P . C. Woodland, and J. Such. Case-bench: Context-aware safety benchmark for large language models, 2025. URLhttps://arxiv.org/abs/2501.14940
2025 arXiv
-
[84]
J. Sun, Q. Zhang, Y. Duan, X. Jiang, C. Cheng, and R. Xu. Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning, 2024. URL https://arxiv.org/abs/2309. 11359
2024
-
[85]
J. Spataro. Introducing Microsoft 365 Copilot — your copilot for work. https://blogs.microsoft. com/blog/2023/03/16/introducing-microsoft-365-copilot-your-copilot-for-work/ , 2023. Microsoft Official Blog, accessed 22 May 2025
2023
-
[86]
Wallace, S
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh. Universal adversarial triggers for attacking and analyzing NLP. In K. Inui, J. Jiang, V . Ng, and X. Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inter...
2019 doi
-
[87]
X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig. Openhands: An open platform for ai software develo...
2024
-
[88]
Trivedi, T
H. Trivedi, T. Khot, M. Hartmann, R. R. Manku, V . Dong, E. Li, S. Gupta, A. Sabharwal, and N. Balasubramanian. Appworld: A controllable world of apps and people for benchmarking interactive coding agents.ArXiv, abs/2407.18901, 2024. URL https://api.semanticscholar. org/Corpus...
2024 arXiv
-
[89]
Z. Wang, H. Li, R. Zhang, Y. Liu, W. Jiang, W. Fan, Q. Zhao, and G. Xu. Mpma: Preference manipulation attack against model context protocol, 2025. URL https://arxiv.org/abs/2505. 11154
2025
-
[90]
Z. Wang, V . Siu, Z. Ye, T. Shi, Y. Nie, X. Zhao, C. Wang, W. Guo, and D. Song. Agentfuzzer: Generic black-box fuzzing for indirect prompt injection against llm agents, 2025. URL https: //arxiv.org/abs/2505.05849
2025 arXiv
-
[91]
Z. Wang, T. Pang, C. Du, M. Lin, W. Liu, and S. Yan. Better diffusion models further improve adversarial training, 2023. URLhttps://arxiv.org/abs/2302.04638
2023 arXiv
-
[92]
S. Wu, S. Zhao, Q. Huang, K. Huang, M. Yasunaga, K. Cao, V . N. Ioannidis, K. Subbian, J. Leskovec, and J. Zou. Avatar: Optimizing llm agents for tool usage via contrastive reasoning, 2024. URL https://arxiv.org/abs/2406.11200
2024 arXiv
-
[93]
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V . Zhong, and T. Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.ArXiv, abs/2404.07972, 2024....
2024 arXiv
-
[94]
Werner, K
J. Werner, K. Chu, C. Weber, and S. Wermter. Llm-based interactive imitation learning for robotic manipulation, 2025. URLhttps://arxiv.org/abs/2504.21769
2025 arXiv
-
[95]
Y. Yao, X. Tong, R. Wang, Y. Wang, L. Li, L. Liu, Y. Teng, and Y. Wang. A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos, 2025. URL https://arxiv.org/abs/ 2502.15806
2025 arXiv
-
[96]
F. Yin, P . Laban, X. Peng, Y. Zhou, Y. Mao, V . Vats, L. Ross, D. Agarwal, C. Xiong, and C.-S. Wu. Bingoguard: Llm content moderation tools with risk levels.ArXiv, abs/2503.06550, 2025. URL https://api.semanticscholar.org/CorpusID:276903386
2025 arXiv
-
[97]
T. Xie, X. Qi, Y. Zeng, Y. Huang, U. M. Sehwag, K. Huang, L. He, B. Wei, D. Li, Y. Sheng, R. Jia, B. Li, K. Li, D. Chen, P . Henderson, and P . Mittal. Sorry-bench: Systematically evaluating large language model safety refusal. InThe Thirteenth International Conference on Lear...
-
[98]
Y. Yuan, W. Jiao, W. Wang, J. tse Huang, J. Xu, T. Liang, P . He, and Z. Tu. Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training, 2024. URL https: //arxiv.org/abs/2407.09121
2024 arXiv
-
[99]
W. Zeng, Y. Liu, R. Mullins, L. Peran, J. Fernandez, H. Harkous, K. Narasimhan, D. Proud, P . Kumar, B. Radharapu, O. Sturman, and O. Wahltinez. Shieldgemma: Generative ai content moderation based on gemma, 2024. URLhttps://arxiv.org/abs/2407.21772
2024 arXiv
-
[100]
Y. Zeng, K. Klyman, A. Zhou, Y. Yang, M. Pan, R. Jia, D. Song, P . Liang, and B. Li. Ai risk categorization decoded (air 2024): From government regulations to corporate policies, 2024. URL https://arxiv.org/abs/2406.17864
2024 arXiv
-
[101]
K. You, H. Zhang, E. Schoop, F. Weers, A. Swearngin, J. Nichols, Y. Yang, and Z. Gan. Ferret-ui: Grounded mobile ui understanding with multimodal llms. InEuropean Conference on Computer Vision, 2024. URLhttps://api.semanticscholar.org/CorpusID:269005503. 17
2024
-
[102]
Y. Zeng, Y. Yang, A. Zhou, J. Z. Tan, Y. Tu, Y. Mai, K. Klyman, M. Pan, R. Jia, D. Song, P . Liang, and B. Li. AIR-BENCH 2024: A safety benchmark based on regulation and policies specified risk categories. InThe Thirteenth International Conference on Learning Representations, ...
2024
-
[103]
Zhang, T
J. Zhang, T. Lan, M. Zhu, Z. Liu, T. Hoang, S. Kokane, W. Yao, J. Tan, A. Prabhakar, H. Chen, Z. Liu, Y. Feng, T. Awalgaonkar, R. Murthy, E. Hu, Z. Chen, R. Xu, J. C. Niebles, S. Heinecke, H. Wang, S. Savarese, and C. Xiong. xlam: A family of large action models to empower ai ...
2024 arXiv
-
[104]
Zhang, T
Y. Zhang, T. Yu, and D. Yang. Attacking vision-language computer agents via pop-ups, 2025. URL https://arxiv.org/abs/2411.02391
2025 arXiv
-
[105]
Y. Zeng, Y. Wu, X. Zhang, H. Wang, and Q. Wu. Autodefense: Multi-agent llm defense against jail- break attacks.ArXiv, abs/2403.04783, 2024. URL https://api.semanticscholar.org/CorpusID: 268297202
2024 arXiv
-
[106]
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023. URLhttps://arxiv.org/abs/2307.15043
2023 arXiv
-
[107]
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks. Improving alignment and robustness with circuit breakers, 2024. URL https: //arxiv.org/abs/2406.04313. 18
2024 arXiv
-
[109]
Zheng, B
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su. Gpt-4v(ision) is a generalist web agent, if grounded. ArXiv, abs/2401.01614, 2024. URLhttps://api.semanticscholar.org/CorpusID:266741821
2024 arXiv
-
[2023]
URLhttps://arxiv.org/abs/2312.06674
-
[2024]
URLhttps://arxiv.org/abs/2410.13886
-
[2025]
URLhttps://openreview.net/forum?id=YfKNaRktan
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.