Pith. sign in

REVIEW 5 cited by

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18693 v1 pith:QW5O4CD2 submitted 2024-12-24 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords attacksdiversegoalsgenerateautomatedeffectivegeneratingmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated red teaming can discover rare model failures and generate challenging examples that can be used for training or evaluation. However, a core challenge in automated red teaming is ensuring that the attacks are both diverse and effective. Prior methods typically succeed in optimizing either for diversity or for effectiveness, but rarely both. In this paper, we provide methods that enable automated red teaming to generate a large number of diverse and successful attacks. Our approach decomposes the task into two steps: (1) automated methods for generating diverse attack goals and (2) generating effective attacks for those goals. While we provide multiple straightforward methods for generating diverse goals, our key contributions are to train an RL attacker that both follows those goals and generates diverse attacks for those goals. First, we demonstrate that it is easy to use a large language model (LLM) to generate diverse attacker goals with per-goal prompts and rewards, including rule-based rewards (RBRs) to grade whether the attacks are successful for the particular goal. Second, we demonstrate how training the attacker model with multi-step RL, where the model is rewarded for generating attacks that are different from past attempts further increases diversity while remaining effective. We use our approach to generate both prompt injection attacks and prompts that elicit unsafe responses. In both cases, we find that our approach is able to generate highly-effective and considerably more diverse attacks than past general red-teaming approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Jailbreaking to Jailbreak

    cs.CL 2025-02 conditional novelty 7.0 of 10

    A transferable multi-turn jailbreak turns refusal-trained black-box LLMs into willing automated jailbreakers, with high attack success against other models and against themselves.

  2. GPT-Red: Automated Red Teaming via Self-Play at Scale

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A self-play-trained red-teaming agent, GPT-Red, discovers prompt injection attacks and is used to adversarially harden GPT-5.6, cutting attack success rates to near zero on several benchmarks.

  3. From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Seed2Harvest expands 1,000 human adversarial prompts into 27,650 LLM-generated variants that keep roughly comparable unsafe-image trigger rates and add hundreds of new geographic contexts.

  4. Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    QDRT combines behavior-conditioned RL, multiple specialized attackers, and a MAP-Elites replay buffer to generate LLM attacks that are more toxic and cover more risk-category/style combinations.

  5. Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning

    cs.AI 2025-06 reject novelty 6.0 of 10

    A three-stage RL framework (cold start, diversity warm-up, curriculum jailbreak) trains a 7B red-team model that reports SOTA jailbreak ASR and diversity on HarmBench, though the evaluation is compromised by training-...

Pith tools