Pith. sign in

REVIEW 2 major objections 3 minor 1 cited by

A Red Teaming Roadmap Towards System-Level Safety

T0 review · 2 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read LLM red teaming should target product safety specs and system-level defenses, not abstract harms.

desk verdict A clear, well-organized position paper that reprioritizes red-teaming research; its central empirical assumption about where real-world harms live is undefended, but the roadmap is still worth engaging. read the letter →

arxiv 2506.05376 v2 pith:Q6C3I4FI submitted 2025-05-30 cs.CR cs.AI

classification cs.CRcs.AI
keywords LLMredteamingproductsafetyspecificationsrealisticthreatmodelssystem-leveljailbreakingsafeguardevaluationAIdeploymentmonitoringmulti-turnattacks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This is a position paper about how AI red teaming research should be spent. The authors argue that the field's current emphasis on jailbreaking models in isolation to test whether they refuse abstractly harmful requests is misdirected. They propose two priorities: red teaming should test violations of a specific product's safety specification under realistic threat models, and it should move from model-level robustness to system-level safety, where monitoring, user banning, and rapid patching catch harm that individual refusals miss. If they are right, red teaming research would shift from benchmark jailbreaks toward deployment-aware evaluation and defense, which they argue is the only way to keep up with chatbots, audio assistants, video generators, and autonomous agents.

What carries the argument

The load-bearing mechanism is the three-level distinction among model, product, and system, together with the product safety specification as the red teaming objective. A product safety specification states what behavior is prohibited for that specific deployment, including who the users are, what tools are attached, and what regulations apply, and the paper treats it as the target that a successful red team finding must violate. The second mechanism is the system-level mitigation loop: trajectory monitoring classifies harm over whole interaction histories, user monitoring detects and bans repeat malicious users, and rapid response patches newly found failures before they are exploited at scale. Red teaming the monitor, through sabotage experiments where an adversary must complete a harmful task without detection, stress-tests that loop. These mechanisms convert the vague question of whether a model is safe into testable questions about a specific product and its deployed defenses.

What would settle it

A head-to-head deployment study would settle it: take the same underlying model in the same product, run one arm with only refusal training and one arm whose red teaming targets the product safety specification with system-level monitoring and user banning, and compare measured harm per active user, for example substantiated abuse reports per 100,000 interactions. If the model-only arm shows equal or lower harm, the paper's priority claim is false.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLM red teaming will reduce real-world harm only if it is re-anchored around two priorities: product safety specifications over abstract social biases or ethical principles, and system-level safety over model-level robustness. It defines a model as the neural network, a product as the deployed application built on it, and a system as the product plus its deployment infrastructure, including monitors, staff, users, and environment, and argues that the attack surface users actually face is the end-to-end product stack, not the bare model. Because safety is contextual, a behavior that violates one product's policy may be acceptable in another, so red teaming objectives should be explicit, actionable, and measured against the target product's stated spec. The paper then argues that realistic threat models differ by product type, with multi-turn conversations for chatbots, prosodic and multilingual channels for audio assistants, frame-spanning harm for video generators, and tool- and environment-driven attacks for agents, and that safeguards should be tested both independently and in realistic sandboxes. Its final step is the claim that system-level measures such as trajectory monitoring, user monitoring, rapid response to newly discovered jailbreaks, and red teaming the monitor itself through sabotage experiments are necessary to make red teaming relevant to deployed risk.

Load-bearing premise

The roadmap assumes that a product safety specification can be written precisely enough to serve as a red teaming objective, and that red teaming against that spec plus system-level mitigation reduces real-world harm more effectively than model-level testing.

Editorial extensions

If this is right

  • Red teaming benchmarks would be evaluated against a stated product safety specification and deployment context, so a jailbreak that violates no deployed product's policy would count as low priority.
  • Multi-turn and trajectory-level attacks would receive more research attention than single-turn jailbreaks, because defenses trained for single turns do not generalize to extended conversations.
  • Safety investments would shift toward asynchronous monitoring, malicious-user detection, and rapid safeguard deployment, since most harm is diffuse over many requests rather than a single catastrophic output.
  • White-box red teaming would assume the adversary can fine-tune the model, and black-box red teaming would assume limited query access, making attacks match real attacker resources.
  • Agent red teaming would require investment in realistic sandboxes and simulated environments, because tool outputs and environmental context are part of the attack surface.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If product-specific specifications become the primary evaluation target, standardized public benchmarks may fragment, because each product's policy is different; a shared harm taxonomy would survive only where regulators impose one.
  • Editorial inference: The emphasis on system-level monitoring implies that model-level alignment research could be deprioritized, yet the system-level loop still depends on the model not producing a single catastrophic output before a monitor can intervene, so the two are complements rather than substitutes.
  • Editorial inference: Red teaming the monitor sets up an arms race between detection evasion and detection improvement; sabotage experiments would need to be repeated continuously as monitors update, making red teaming an ongoing operational process rather than a one-time evaluation.
  • Editorial inference: A testable consequence is that products with user banning and rapid patching should show lower measured abuse per user than products with only refusal-trained models, holding the underlying model fixed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. This position paper argues that LLM red teaming research currently misallocates effort. It recommends three priorities: (1) red teaming should target product safety specifications rather than abstract social biases or ethical principles; (2) threat models should reflect realistic attackers and deployment contexts rather than idealized single-query settings; and (3) red teaming should expand from model-level refusal robustness to system-level safety, including trajectory monitoring, user monitoring, rapid patching, and red teaming of monitors. The paper motivates these claims through a taxonomy of model/product/system, worked threat-model sketches for chatbots, audio assistants, video generators, and agents, and a set of best practices in Section 4. Section 5 develops the system-level argument, and Section 6 considers alternative views. The paper is explicitly a roadmap rather than an empirical study; its central claims are normative and strategic.

Significance. If the roadmap is accepted, it would redirect a substantial portion of red teaming research from universal jailbreak benchmarks toward deployment-aware evaluation and system-level defenses. The paper's contribution is mainly organizational and rhetorical: it clarifies distinctions among model, product, and system; identifies understudied attack surfaces such as multi-turn conversations, audio, video, agents, and tool outputs; and offers a practical best-practices list in Section 4. It is honest about its status as a position paper, explicitly engaging with an alternative view in Section 6. The main weakness is that the two prioritizations—product over abstract, and system over model—are supported by selected examples rather than by systematic evidence, leaving the strongest comparative claims underdetermined. As a position paper, it need not contain new experiments, but its empirical premises should be supported with a structured literature synthesis or measurement study.

major comments (2)
  1. [Section 5, 'Trajectory and User Monitoring' and 'Rapid Response'] The central comparative claim that system-level safety should be prioritized over model-level robustness rests on the empirical premise that 'many realistic harmful uses of an AI system may require its cooperation over multiple turns' and that asynchronous monitoring can catch failures before harm scales. The manuscript provides examples and citations (e.g., [30, 47, 56]) but no evidence about the relative frequency or severity of diffuse multi-turn harms versus single-output catastrophic harms. The introduction's motivating examples—dangerous biological agents and other catastrophic risks—are precisely the kind of single-response harms for which post-hoc user banning and patching cannot undo the first harmful output. The sentence 'as long as sufficiently effective asynchronous monitoring detects these failures as they happen' is a conditional, not a demonstrated premise, and it is hardest for the highest-severity cases. To support the prioritization, the paper should either supply or cite data on harm distributions and intervention lead times in deployed products, or soften the claim from 'should be prioritized over' to 'should complement.'
  2. [Section 2, 'Safety Specifications Should Focus on Product-Level Risk'] The first pillar assumes that a product safety specification can be defined precisely enough to serve as a red teaming objective and that such specifications are tied to actual societal harm. The paper acknowledges that products can and should have divergent safety considerations and that many products lack well-defined specifications; its fallback is that researchers should 'define or infer a plausible policy context.' This fallback reintroduces the same normative ambiguity that the paper criticizes in abstract harm categories, because a plausible policy may be vague or contested. The manuscript does not present a worked example in which a product specification is translated into a concrete red teaming metric (what counts as a successful breach, what the evaluation metric is) for a specific product. Without at least one such case study, the claim that product-level red teaming is more tractable and more decision-relevant than model-level harm testing remains an assertion. A worked example would make the proposal easier to adopt and would make its central premise falsifiable.
minor comments (3)
  1. [Section 4, 'Assessing the Delta'] The sentence 'There would be no difference in the models vulnerabilities' contains a typo; it should read 'the model's vulnerabilities.'
  2. [Section 3.1, first bullet list] The claim that 'more difficult single-turn attacks seem to be independent, ad-hoc issues with safeguards that can be quickly patched' is supported by a single self-cited reference [63]; a systematic comparison or an independent replication would make the argument more robust.
  3. [Section 6, 'Argument for Static Harm Research'] The alternative view is presented fairly but dismissed in a few sentences; readers who disagree with the product-specification premise would benefit from a more detailed rebuttal, particularly concerning the role of regulation and societal consensus.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a position piece whose recommendations rest on contextual arguments and conditional empirical claims, not on self-citations or fitted quantities.

full rationale

This is a position paper with no equations, fitted parameters, or formal derivation chain. Its central recommendation—prioritize product safety specifications and system-level safeguards—is argued from definitions and stated value judgments, not reduced to its own inputs. The paper does cite several works from its own team (e.g., [47] on multi-turn jailbreaks, [63] on rapid response, and [41]/[90] on agent and browser red teaming), but these are supporting examples for empirical premises that are also supported by independent citations or are externally testable; the recommendations do not reduce to those citations. The load-bearing conditional about asynchronous monitoring is explicitly conditional and is not presented as a demonstrated premise. No self-definitional step, fitted input, imported uniqueness theorem, or ansatz-smuggling chain is present. Concerns about unquantified harm distributions are evidence gaps or correctness risks, not circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters or invented entities are present. The paper introduces conceptual distinctions such as model, product, and system, and trajectory-level versus user-level monitoring, but these are definitions rather than new physical or computational entities. The load-bearing content is the set of domain assumptions listed above.

assumptions (5)
  • domain assumption Product safety specifications can be defined precisely enough to serve as measurable red teaming objectives.
    Section 2 assumes red teaming objectives should derive from product specifications; the paper does not demonstrate that such specifications exist and are evaluable across products.
  • domain assumption Threat models that approximate real attackers are more useful for mitigating real-world harm than abstract or idealized models.
    Section 3 opens with this premise; no empirical comparison is provided.
  • domain assumption System-level mitigations such as monitoring, user banning, and rapid patching are available and effective for deployed AI products.
    Section 5 relies on deployer control of the stack; the paper acknowledges open-weight and distributed cases only partially.
  • domain assumption Single-turn jailbreak failures are largely ad-hoc and patchable, while multi-turn and system-level harms are more fundamental.
    Section 3.1 states this distinction with selected citations; it is load-bearing for the call to shift priorities.
  • domain assumption Current conference red teaming submissions do not prioritize the right research problems.
    Section 1 asserts this without systematic meta-analysis; it is the motivational premise of the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Red Teaming Roadmap Towards System-Level Safety." pith.science (2026). https://pith.science/paper/Q6C3I4FI

@misc{pith2026250605376,
  author       = {Pith},
  title        = {Pith review of: A Red Teaming Roadmap Towards System-Level Safety},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q6C3I4FI}},
  note         = {Machine review of arXiv:2506.05376}
}
read the original abstract

Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has effectively identified critical vulnerabilities in state-of-the-art refusal-trained LLMs. However, in our view the many conference submissions on LLM red teaming do not, in aggregate, prioritize the right research problems. First, testing against clear product safety specifications should take a higher priority than abstract social biases or ethical principles. Second, red teaming should prioritize realistic threat models that represent the expanding risk landscape and what real attackers might do. Finally, we contend that system-level safety is a necessary step to move red teaming research forward, as AI models present new threats as well as affordances for threat mitigation (e.g., detection and banning of malicious users) once placed in a deployment context. Adopting these priorities will be necessary in order for red teaming research to adequately address the slate of new threats that rapid AI advances present today and will present in the very near future.

Figures

Figures reproduced from arXiv: 2506.05376 by the authors.

Figure 1
Figure 1. Top: An overview of this paper’s assertion that red teams should prioritize assessments based [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. An illustration of user monitoring, trajectory monitoring and monitor red teaming. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Across 2M+ in-the-wild LLM conversations, jailbreak attempts show no higher complexity than normal chats, and assistant toxicity has declined over time, suggesting bounded attack sophistication.

Reference graph

Works this paper leans on

111 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [1]

    URL https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995

    Anthropic. URL https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995. pdf

  2. [2]

    Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku

    Anthropic_Computer_Use. Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku. URLhttps://www.anthropic.com/news/3-5-models-and-computer-use

  3. [3]

    Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield- Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan. ...

  4. [4]

    Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran- Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. ...

  5. [5]

    Baker, J

    B. Baker, J. Huizinga, L. Gao, Z. Dou, M. Y. Guan, A. Madry, W. Zaremba, J. Pachocki, and D. Farhi. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation, 2025. URL https://arxiv.org/abs/2503.11926

  6. [6]

    B. R. Bartoldson, J. Diffenderfer, K. Parasyris, and B. Kailkhura. Adversarial robustness limits via scaling-law and human-alignment studies.arXiv preprint arXiv:2404.09349, 2024

  7. [7]

    Bengio, G

    Y. Bengio, G. Hinton, A. Yao, D. Song, P . Abbeel, S. Russell, P . Torr, J. Brauner, S. Mindermann, et al. Managing extreme AI risks amid rapid progress.Science, 384(6698):916–919, 2024. doi: 10.1126/science.adn0117

  8. [8]

    Benton, M

    J. Benton, M. Wagner, E. Christiansen, C. Anil, E. Perez, J. Srivastav, E. Durmus, D. Ganguli, S. Kravec, B. Shlegeris, J. Kaplan, H. Karnofsky, E. Hubinger, R. Grosse, S. R. Bowman, and D. Duvenaud. Sabotage evaluations for frontier models, 2024. URL https://arxiv.org/abs/ 2410.21514

Show all 111 references
  1. [9]

    A. N. Bhagoji, W. He, B. Li, and D. Song. Practical black-box attacks on deep neural networks using efficient query mechanisms. InProceedings of the European conference on computer vision (ECCV), pages 154–169, 2018

  2. [10]

    Bonatti, D

    R. Bonatti, D. Zhao, F. Bonacci, D. Dupont, S. Abdali, Y. Li, Y. Lu, J. Wagle, K. Koishida, A. F. C. Bucker, L. Jang, and Z. Hui. Windows agent arena: Evaluating multi-modal os agents at scale. ArXiv, abs/2409.08264, 2024. URLhttps://api.semanticscholar.org/CorpusID:272600411. 11

  3. [11]

    Carlini and D

    N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. In2017 ieee symposium on security and privacy (sp), pages 39–57. Ieee, 2017

  4. [12]

    P . Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong. Jailbreaking black box large language models in twenty queries.arXiv preprint arXiv:2310.08419, 2023. URL https: //arxiv.org/abs/2310.08419

  5. [13]

    C. Chen, Z. Zhang, B. Guo, S. Ma, I. Khalilov, S. A. Gebreegziabher, Y. Ye, Z. Xiao, Y. Yao, T. Li, and T. J.-J. Li. The obvious invisible threat: Llm-powered gui agents’ vulnerability to fine-print injections, 2025. URLhttps://arxiv.org/abs/2504.11281

  6. [14]

    G. Chen, F. Song, Z. Zhao, X. Jia, Y. Liu, Y. Qiao, and W. Zhang. Audiojailbreak: Jailbreak attacks against end-to-end large audio-language models, 2025. URL https://arxiv.org/abs/ 2505.14103

  7. [15]

    P .-Y. Chen, H. Zhang, Y. Sharma, J. Yi, and C.-J. Hsieh. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. InProceedings of the 10th ACM workshop on artificial intelligence and security, pages 15–26, 2017

  8. [16]

    Cheng, C

    G. Cheng, C. Zhang, W. Cai, L. Zhao, C. Sun, and J. Bian. Empowering large language models on robotic manipulation with affordance prompting, 2024. URL https://arxiv.org/abs/2404. 11027

  9. [17]

    Croce and M

    F. Croce and M. Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. InICML, 2020

  10. [18]

    Croce, M

    F. Croce, M. Andriushchenko, V . Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P . Mittal, and M. Hein. Robustbench: a standardized adversarial robustness benchmark.arXiv preprint arXiv:2010.09670, 2020

  11. [19]

    G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu. MASTERKEY: Automated jailbreaking of large language model chatbots. InProceedings of the 2024 Network and Distributed System Security Symposium (NDSS), 2024

  12. [20]

    X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su. Mind2web: Towards a generalist agent for the web.ArXiv, abs/2306.06070, 2023. URL https://api.semanticscholar. org/CorpusID:259129428

  13. [21]

    For-profit AI safety: AI safety needs to scale and here’s how you can do it

    Esben Kran. For-profit AI safety: AI safety needs to scale and here’s how you can do it. https:// apartresearch.com/news/ai-safety-needs-to-scale-and-heres-how-you-can-do-it , 2024. Apart Research blog, accessed 22 May 2025

  14. [22]

    P . M. Gade, S. Lermen, C. Rogers-Smith, and J. Ladish. Badllama: cheaply removing safety fine- tuning from llama 2-chat 13b.ArXiv, abs/2311.00117, 2023. URL https://api.semanticscholar. org/CorpusID:264832925

  15. [23]

    Ghosh, P

    S. Ghosh, P . Varshney, E. Galinkin, and C. Parisien. Aegis: Online adaptive ai content safety moderation with ensemble of llm experts.ArXiv, abs/2404.05993, 2024. URL https://api. semanticscholar.org/CorpusID:269009460

  16. [24]

    I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples.arXiv preprint arXiv:1412.6572, 2014

  17. [25]

    Google. Veo. URLhttps://deepmind.google/models/veo/. 12

  18. [26]

    Greenblatt, B

    R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger. Ai control: Improving safety despite intentional subversion, 2024. URLhttps://arxiv.org/abs/2312.06942

  19. [27]

    M. Y. Guan, M. Joglekar, E. Wallace, S. Jain, B. Barak, A. Helyar, R. Dias, A. Vallone, H. Ren, J. Wei, H. W. Chung, S. Toyer, J. Heidecke, A. Beutel, and A. Glaese. Deliberative alignment: Reasoning enables safer language models, 2025. URLhttps://arxiv.org/abs/2412.16339

  20. [28]

    Z. Guan, M. Hu, R. Zhu, S. Li, and A. Vullikanti. Benign samples matter! fine-tuning on outlier benign samples severely breaks safety. 2025. URL https://api.semanticscholar.org/CorpusID: 278501392

  21. [29]

    Hafez, A

    A. Hafez, A. N. Akhormeh, A. Hegazy, and A. Alanwar. Safe llm-controlled robots with formal guarantees via reachability analysis, 2025. URLhttps://arxiv.org/abs/2503.03911

  22. [30]

    Automated multi-turn red-teaming with cascade

    Haize. Automated multi-turn red-teaming with cascade. URL https://www.haizelabs.com/ technology/automated-multi-turn-red-teaming-with-cascade

  23. [31]

    Hammond, A

    L. Hammond, A. Chan, J. Clifton, J. Hoelscher-Obermaier, A. Khan, E. McLean, C. Smith, W. Bar- fuss, J. Foerster, T. Gavenˇ ciak, T. A. Han, E. Hughes, V . Kovaˇ rík, J. Kulveit, J. Z. Leibo, C. Oester- held, C. S. de Witt, N. Shah, M. Wellman, P . Bova, T. Cimpeanu, C. Ezell,...

  24. [32]

    W. Held, M. J. Ryan, A. Shrivastava, A. S. Khan, C. Ziems, E. Li, M. Bartelds, M. Sun, T. Li, W. Gan, and D. Yang. Cava: Comprehensive assessment of voice assistants. https://github. com/SALT-NLP/CAVA, 2025. URL https://talkarena.org/cava. A benchmark for evaluating large audi...

  25. [33]

    Hendrycks, M

    D. Hendrycks, M. Mazeika, and T. Woodside. An overview of catastrophic ai risks, 2023. URL https://arxiv.org/abs/2306.12001

  26. [34]

    Huang, S

    T. Huang, S. Hu, F. Ilhan, S. F. Tekin, and L. Liu. Harmful fine-tuning attacks and defenses for large language models: A survey, 2024. URLhttps://arxiv.org/abs/2409.18169

  27. [35]

    Hughes, S

    J. Hughes, S. Price, A. Lynch, R. Schaeffer, F. Barez, S. Koyejo, H. Sleight, E. Jones, E. Perez, and M. Sharma. Best-of-n jailbreaking, 2024. URLhttps://arxiv.org/abs/2412.03556

  28. [36]

    H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa. Llama guard: Llm-based input-output safeguard for human-ai conversations,

  29. [37]

    Jones, A

    E. Jones, A. Dragan, and J. Steinhardt. Adversaries can misuse combinations of safe models, 2024. URLhttps://arxiv.org/abs/2406.14595

  30. [38]

    J. Y. Koh, R. Lo, L. Jang, V . Duvvur, M. C. Lim, P .-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.ArXiv, abs/2401.13649, 2024. URLhttps://api.semanticscholar.org/CorpusID:267199749

  31. [39]

    Kritz, V

    J. Kritz, V . Robinson, R. Vacareanu, B. Varjavand, M. Choi, B. Gogov, S. R. Team, S. Yue, W. E. Primack, and Z. Wang. Jailbreaking to jailbreak, 2025. URL https://arxiv.org/abs/2502.09638

  32. [40]

    Krizhevsky, V

    A. Krizhevsky, V . Nair, and G. Hinton. Cifar-10 (canadian institute for advanced research). URL http://www.cs.toronto.edu/~kriz/cifar.html. 13

  33. [41]

    Kumar, E

    P . Kumar, E. Lau, S. Vijayakumar, T. Trinh, S. R. Team, E. Chang, V . Robinson, S. Hendryx, S. Zhou, M. Fredrikson, S. Yue, and Z. Wang. Refusal-trained llms are easily jailbroken as browser agents,

  34. [42]

    M. Kuo, J. Zhang, A. Ding, Q. Wang, L. DiValentin, Y. Bao, W. Wei, H. Li, and Y. Chen. H- cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash thinking, 2025. URL https: //...

  35. [43]

    Laban, H

    P . Laban, H. Hayashi, Y. Zhou, and J. Neville. Llms get lost in multi-turn conversation, 2025. URL https://arxiv.org/abs/2505.06120

  36. [44]

    H. Lai, X. Liu, I. L. Iong, S. Yao, Y. Chen, P . Shen, H. Yu, H. Zhang, X. Zhang, Y. Dong, and J. Tang. Autowebglm: A large language model-based web navigating agent.Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024. URL https: //api.se...

  37. [45]

    Le Jeune, J

    P . Le Jeune, J. Liu, L. Rossi, and M. Dora. Realharm: A collection of real-world language model application failures.arXiv preprint arXiv:2504.10277, 2025

  38. [46]

    A. Li, Y. Zhou, V . C. Raghuram, T. Goldstein, and M. Goldblum. Commercial llm agents are already vulnerable to simple yet dangerous attacks, 2025. URLhttps://arxiv.org/abs/2502.08586

  39. [47]

    N. Li, Z. Han, I. Steneker, W. Primack, R. Goodside, H. Zhang, Z. Wang, C. Menghini, and S. Yue. Llm defenses are not robust to multi-turn human jailbreaks yet, 2024. URL https: //arxiv.org/abs/2408.15221

  40. [48]

    J. Liu, S. Liang, S. Zhao, R. Tu, W. Zhou, X. Cao, D. Tao, and S. K. Lam. Jailbreaking the text-to-video generative models, 2025. URLhttps://arxiv.org/abs/2505.06679

  41. [49]

    X. Liu, Z. Yu, Y. Zhang, N. Zhang, and C. Xiao. Automatic and universal prompt injection attacks against large language models.arXiv preprint arXiv:2403.04957, 2024

  42. [50]

    Y. Lu, T. Ju, M. Zhao, X. Ma, Y. Guo, and Z. Zhang. Eva: Red-teaming gui agents via evolving indirect prompt injection, 2025. URLhttps://arxiv.org/abs/2505.14289

  43. [51]

    Madry, A

    A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks.arXiv preprint arXiv:1706.06083, 2017

  44. [52]

    Mazeika, L

    M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024

  45. [53]

    A. MCP . Introducing the model context protocol. URL https://www.anthropic.com/news/ model-context-protocol

  46. [54]

    Mehrotra, M

    A. Mehrotra, M. Zampetakis, P . Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi. Tree of attacks: Jailbreaking black-box llms automatically.arXiv preprint arXiv:2312.02119, 2023. URL https://arxiv.org/abs/2312.02119

  47. [55]

    Y. Miao, Y. Zhu, Y. Dong, L. Yu, J. Zhu, and X.-S. Gao. T2vsafetybench: Evaluating the safety of text-to-video generative models, 2024. URLhttps://arxiv.org/abs/2407.05965

  48. [56]

    Nakash, G

    I. Nakash, G. Kour, G. Uziel, and A. Anaby-Tavor. Breaking react agents: Foot-in-the-door attack will get you in, 2024. URLhttps://arxiv.org/abs/2410.16950. 14

  49. [57]

    Narodytska and S

    N. Narodytska and S. P . Kasiviswanathan. Simple black-box adversarial perturbations for deep networks.arXiv preprint arXiv:1612.06299, 2016

  50. [58]

    Nguyen, S

    V .-A. Nguyen, S. Zhao, G. Dao, R. Hu, Y. Xie, and L. A. Tuan. Three minds, one legend: Jailbreak large reasoning model with adaptive stacked ciphers, 2025. URL https://arxiv.org/abs/2505. 16241

  51. [59]

    Sora: Creating video from text

    OpenAI. Sora: Creating video from text. https://openai.com/index/sora/, 2024. Ac- cessed 22 May 2025

  52. [60]

    URLhttps://cdn.openai.com/operator_system_card.pdf

    Operator. URLhttps://cdn.openai.com/operator_system_card.pdf

  53. [61]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P . Welinder, P . Christiano, J. Leike, and R. Lowe. Training language models to follow instructio...

  54. [62]

    Papernot, P

    N. Papernot, P . McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami. Practical black-box attacks against machine learning. InProceedings of the 2017 ACM on Asia conference on computer and communications security, pages 506–519, 2017

  55. [63]

    A. Peng, J. Michael, H. Sleight, E. Perez, and M. Sharma. Rapid response: Mitigating llm jailbreaks with a few examples, 2024. URLhttps://arxiv.org/abs/2411.07494

  56. [64]

    Perez, S

    E. Perez, S. Ringer, K. Lukoši¯ut˙e, K. Nguyen, C. Pettit, C. Olsson, and et al. Discovering language model behaviors with model-written evaluations. InFindings of ACL 2023, pages 3419–3448, 2023

  57. [65]

    S. Pichai. Cloud next 2023: Sharing the best of our AI with google cloud. https://blog.google/ products/google-cloud/cloud-next-2023-sundar-pichai-keynote/ , 2023. Google Blog, ac- cessed 22 May 2025

  58. [66]

    X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P . Mittal, and P . Henderson. Safety alignment should be made more than just a few tokens deep, 2024. URL https://arxiv.org/abs/2406. 05946

  59. [67]

    X. Qi, Y. Zeng, T. Xie, P .-Y. Chen, R. Jia, P . Mittal, and P . Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=hTEGyKf0dZ

  60. [68]

    Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y.-T. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, R. Tian, R. Xie, J. Zhou, M. H. Gerstein, D. Li, Z. Liu, and M. Sun. Toolllm: Facilitating large language models to master 16000+ real-world apis.ArXiv, abs/2307.16789, 2023. URL htt...

  61. [69]

    Radosevich and J

    B. Radosevich and J. Halloran. Mcp safety audit: Llms with the model context protocol allow major security exploits, 2025. URLhttps://arxiv.org/abs/2504.03767

  62. [70]

    Rando, J

    J. Rando, J. Zhang, N. Carlini, and F. Tramèr. Adversarial ml problems are getting harder to solve and to evaluate, 2025. URLhttps://arxiv.org/abs/2502.02260

  63. [71]

    M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, R. Comanescu, C. Akbulut, T. Stepleton, J. Mateos-Garcia, S. Bergman, J. Kay, C. Griffin, B. Bariach, I. Gabriel, V . Rieser, W. Isaac, and L. Weidinger. Gaps in the safety evaluation of generative ai. InProceedings of the 2024...

  64. [72]

    Q. Ren, H. Li, D. Liu, Z. Xie, X. Lu, Y. Qiao, L. Sha, J. Yan, L. Ma, and J. Shao. Derail yourself: Multi-turn llm jailbreak attack through self-discovered clues, 2024. URL https://arxiv.org/abs/ 2410.10700

  65. [73]

    Robey, Z

    A. Robey, Z. Ravichandran, V . Kumar, H. Hassani, and G. J. Pappas. Jailbreaking llm-controlled robots.arXiv preprint arXiv:2410.13691, 2024

  66. [74]

    J. Roh, V . Shejwalkar, and A. Houmansadr. Multilingual and multi-accent jailbreaking of audio llms, 2025. URLhttps://arxiv.org/abs/2504.01094

  67. [75]

    Russinovich, A

    M. Russinovich, A. Salem, and R. Eldan. Great, now write an article about that: The crescendo multi-turn llm jailbreak attack.arXiv preprint arXiv:2404.01833, 2024. URL https://arxiv.org/ abs/2404.01833

  68. [76]

    Sharma, M

    M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, E. Perez, and et al. Towards understanding sycophancy in language models. InProceedings of ICLR 2024, 2024

  69. [77]

    Sharma, M

    M. Sharma, M. Tong, J. Mu, J. Wei, J. Kruthoff, S. Goodfriend, E. Ong, A. Peng, R. Agarwal, C. Anil, A. Askell, N. Bailey, J. Benton, E. Bluemke, S. R. Bowman, E. Christiansen, H. Cunningham, A. Dau, A. Gopal, R. Gilson, L. Graham, L. Howard, N. Kalra, T. Lee, K. Lin, P . Lofg...

  70. [78]

    Sheshadri, A

    A. Sheshadri, A. Ewart, P . Guo, A. Lynch, C. Wu, V . Hebbar, H. Sleight, A. C. Stickland, E. Perez, D. Hadfield-Menell, and S. Casper. Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms.arXiv preprint arXiv:2407.15549, 2024

  71. [79]

    Sikorski, L

    P . Sikorski, L. Schrader, K. Yu, L. Billadeau, J. Meenakshi, N. Mutharasan, F. Esposito, H. AliAkbar- pour, and M. Babaiasl. Deployment of large language models to control mobile robots at the edge,

  72. [80]

    Y. Song, F. Xu, S. Zhou, and G. Neubig. Beyond browsing: Api-based web agents, 2025. URL https://arxiv.org/abs/2410.16464

  73. [81]

    Souly, Q

    A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P . Abbeel, J. Svegliato, S. Emmons, O. Watkins, and S. Toyer. A strongreject for empty jailbreaks, 2024

  74. [82]

    URLhttps://arxiv.org/abs/2405.17670

  75. [83]

    G. Sun, X. Zhan, S. Feng, P . C. Woodland, and J. Such. Case-bench: Context-aware safety benchmark for large language models, 2025. URLhttps://arxiv.org/abs/2501.14940

  76. [84]

    J. Sun, Q. Zhang, Y. Duan, X. Jiang, C. Cheng, and R. Xu. Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning, 2024. URL https://arxiv.org/abs/2309. 11359

  77. [85]

    J. Spataro. Introducing Microsoft 365 Copilot — your copilot for work. https://blogs.microsoft. com/blog/2023/03/16/introducing-microsoft-365-copilot-your-copilot-for-work/ , 2023. Microsoft Official Blog, accessed 22 May 2025

  78. [86]

    Wallace, S

    E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh. Universal adversarial triggers for attacking and analyzing NLP. In K. Inui, J. Jiang, V . Ng, and X. Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inter...

  79. [87]

    X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig. Openhands: An open platform for ai software develo...

  80. [88]

    Trivedi, T

    H. Trivedi, T. Khot, M. Hartmann, R. R. Manku, V . Dong, E. Li, S. Gupta, A. Sabharwal, and N. Balasubramanian. Appworld: A controllable world of apps and people for benchmarking interactive coding agents.ArXiv, abs/2407.18901, 2024. URL https://api.semanticscholar. org/Corpus...

  81. [89]

    Z. Wang, H. Li, R. Zhang, Y. Liu, W. Jiang, W. Fan, Q. Zhao, and G. Xu. Mpma: Preference manipulation attack against model context protocol, 2025. URL https://arxiv.org/abs/2505. 11154

  82. [90]

    Z. Wang, V . Siu, Z. Ye, T. Shi, Y. Nie, X. Zhao, C. Wang, W. Guo, and D. Song. Agentfuzzer: Generic black-box fuzzing for indirect prompt injection against llm agents, 2025. URL https: //arxiv.org/abs/2505.05849

  83. [91]

    Z. Wang, T. Pang, C. Du, M. Lin, W. Liu, and S. Yan. Better diffusion models further improve adversarial training, 2023. URLhttps://arxiv.org/abs/2302.04638

  84. [92]

    S. Wu, S. Zhao, Q. Huang, K. Huang, M. Yasunaga, K. Cao, V . N. Ioannidis, K. Subbian, J. Leskovec, and J. Zou. Avatar: Optimizing llm agents for tool usage via contrastive reasoning, 2024. URL https://arxiv.org/abs/2406.11200

  85. [93]

    T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V . Zhong, and T. Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.ArXiv, abs/2404.07972, 2024....

  86. [94]

    Werner, K

    J. Werner, K. Chu, C. Weber, and S. Wermter. Llm-based interactive imitation learning for robotic manipulation, 2025. URLhttps://arxiv.org/abs/2504.21769

  87. [95]

    Y. Yao, X. Tong, R. Wang, Y. Wang, L. Li, L. Liu, Y. Teng, and Y. Wang. A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos, 2025. URL https://arxiv.org/abs/ 2502.15806

  88. [96]

    F. Yin, P . Laban, X. Peng, Y. Zhou, Y. Mao, V . Vats, L. Ross, D. Agarwal, C. Xiong, and C.-S. Wu. Bingoguard: Llm content moderation tools with risk levels.ArXiv, abs/2503.06550, 2025. URL https://api.semanticscholar.org/CorpusID:276903386

  89. [97]

    T. Xie, X. Qi, Y. Zeng, Y. Huang, U. M. Sehwag, K. Huang, L. He, B. Wei, D. Li, Y. Sheng, R. Jia, B. Li, K. Li, D. Chen, P . Henderson, and P . Mittal. Sorry-bench: Systematically evaluating large language model safety refusal. InThe Thirteenth International Conference on Lear...

  90. [98]

    Y. Yuan, W. Jiao, W. Wang, J. tse Huang, J. Xu, T. Liang, P . He, and Z. Tu. Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training, 2024. URL https: //arxiv.org/abs/2407.09121

  91. [99]

    W. Zeng, Y. Liu, R. Mullins, L. Peran, J. Fernandez, H. Harkous, K. Narasimhan, D. Proud, P . Kumar, B. Radharapu, O. Sturman, and O. Wahltinez. Shieldgemma: Generative ai content moderation based on gemma, 2024. URLhttps://arxiv.org/abs/2407.21772

  92. [100]

    Y. Zeng, K. Klyman, A. Zhou, Y. Yang, M. Pan, R. Jia, D. Song, P . Liang, and B. Li. Ai risk categorization decoded (air 2024): From government regulations to corporate policies, 2024. URL https://arxiv.org/abs/2406.17864

  93. [101]

    K. You, H. Zhang, E. Schoop, F. Weers, A. Swearngin, J. Nichols, Y. Yang, and Z. Gan. Ferret-ui: Grounded mobile ui understanding with multimodal llms. InEuropean Conference on Computer Vision, 2024. URLhttps://api.semanticscholar.org/CorpusID:269005503. 17

  94. [102]

    Y. Zeng, Y. Yang, A. Zhou, J. Z. Tan, Y. Tu, Y. Mai, K. Klyman, M. Pan, R. Jia, D. Song, P . Liang, and B. Li. AIR-BENCH 2024: A safety benchmark based on regulation and policies specified risk categories. InThe Thirteenth International Conference on Learning Representations, ...

  95. [103]

    Zhang, T

    J. Zhang, T. Lan, M. Zhu, Z. Liu, T. Hoang, S. Kokane, W. Yao, J. Tan, A. Prabhakar, H. Chen, Z. Liu, Y. Feng, T. Awalgaonkar, R. Murthy, E. Hu, Z. Chen, R. Xu, J. C. Niebles, S. Heinecke, H. Wang, S. Savarese, and C. Xiong. xlam: A family of large action models to empower ai ...

  96. [104]

    Zhang, T

    Y. Zhang, T. Yu, and D. Yang. Attacking vision-language computer agents via pop-ups, 2025. URL https://arxiv.org/abs/2411.02391

  97. [105]

    Y. Zeng, Y. Wu, X. Zhang, H. Wang, and Q. Wu. Autodefense: Multi-agent llm defense against jail- break attacks.ArXiv, abs/2403.04783, 2024. URL https://api.semanticscholar.org/CorpusID: 268297202

  98. [106]

    A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023. URLhttps://arxiv.org/abs/2307.15043

  99. [107]

    A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks. Improving alignment and robustness with circuit breakers, 2024. URL https: //arxiv.org/abs/2406.04313. 18

  100. [109]

    Zheng, B

    B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su. Gpt-4v(ision) is a generalist web agent, if grounded. ArXiv, abs/2401.01614, 2024. URLhttps://api.semanticscholar.org/CorpusID:266741821

  101. [2023]

    URLhttps://arxiv.org/abs/2312.06674

  102. [2024]

    URLhttps://arxiv.org/abs/2410.13886

  103. [2025]

    URLhttps://openreview.net/forum?id=YfKNaRktan

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.