Pith. sign in

REVIEW 4 major objections 6 minor 41 references

A phase-structured multi-turn attack reaches 97–100% success on frontier LLMs, but only 67–79% when full operational detail is required.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 06:39 UTC pith:Z27HHRVP

load-bearing objection Solid methods paper: explicit multi-phase attack program plus dual actionability ASR and a clean depth ablation; numbers are real but upper-bound under judge coupling. the 4 major comments →

arxiv 2607.11151 v1 pith:Z27HHRVP submitted 2026-07-13 cs.CR cs.AI

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

classification cs.CR cs.AI
keywords LLM safetyred-teamingmulti-turn jailbreakadversarial evaluationactionability-gated ASRphase-structured attackmulti-role jury
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Mainstream LLM safety tests still rely on single-turn prompts and a single judge score. This paper argues that those tests systematically understate risk from adaptive multi-turn adversaries and hide how much of a “success” is actually usable. It introduces AMT-X, which turns the attack into an explicit multi-phase state machine (reconnaissance → boundary probing → contradiction → reframing → extraction) driven by graded semantic signals from the victim, and grades success with a multi-role jury whose phase-conditioned checklists separate partial answers from complete, real, operational procedures. On a 6×7 grid of frontier models and Moderation categories, overall success hits 97.6–100%, yet full actionable success falls to 66.7–78.6%—a gap of up to 33 points. A controlled depth ablation shows that capping the attacker before reframing collapses overall success to roughly 29–36%, recovering only when later phases are allowed.

Core claim

When multi-turn jailbreaks are cast as a reproducible phase state machine and scored by checklist-gated dual metrics, frontier models under default alignment yield near-ceiling overall attack success (97.6–100%) but substantially lower fully actionable success (66.7–78.6% mean 71.4%), and that success is strongly depth-dependent: early phases alone produce only ~29–36% overall ASR while reframing-plus-extraction recovers the headline rate.

What carries the argument

AMT-X’s multi-phase state machine (P0–P4) with semantic-analysis-driven transitions, a 31-technique phased library, and dual metrics: lenient overall ASR (score ≥ threshold once any critical actionability item passes) versus stricter full ASR (all critical P4 items: complete structure, real values, and operational steps).

Load-bearing premise

That the frozen multi-role jury’s phase-conditioned checklists—especially the three critical extraction items—faithfully measure actionable harm rather than a scoring signal the attacker is also optimizing against.

What would settle it

Rerun the same 6×7 victim–goal grid under a fully decoupled scorer (or human raters) that the attacker never sees, and check whether full ASR remains in the 67–79% band or collapses once the attacker can no longer steer to the jury’s checklist.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Single headline ASR numbers should be replaced by dual reporting that separates partially vs fully actionable harm.
  • Safety evaluations that stop at single-turn or shallow multi-turn probes will systematically understate risk.
  • Attack success concentrates after reframing and extraction; depth caps or escalation gates after refusal mapping become candidate low-overhead mitigations (suggested, not tested).
  • Reproducible phase machines enable controlled ablations that attribute success to structure rather than free-form improvisation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the 33-point partial-vs-full gap generalizes, many published multi-turn ASR figures may overstate operational risk by counting incomplete procedures as successes.
  • Provider-side turn ceilings or anomaly-triggered session suspension after early refusal mapping could cut extractor-grade success sharply, at the cost of benign long conversations.
  • The same checklist-gated dual metric could be retrofitted to prior multi-turn attacks for head-to-head comparison under a shared actionability standard.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes AMT-X, a multi-turn red-teaming framework that casts adaptive jailbreaking as an explicit P0–P4 phase state machine driven by graded semantic signals (refusal, cooperation, disclosure), a 31-technique phased library with UCB1 selection, and a multi-role debate jury whose phase-conditioned checklists produce dual metrics: a lenient overall ASR (score threshold, one critical item) and a stricter full ASR (all P4 critical items: complete structure, real values, operational steps). On a 6×7 grid of frontier victims × Moderation sub-categories (three stochastic full-pipeline runs), it reports overall ASR 97.6–100% versus full ASR 66.7–78.6% (gap up to 33 pp), and a phase-cap ablation in which overall ASR collapses to ~29–36% at P1/P2 and recovers to 97.6% once reframing/extraction are allowed. The manuscript positions the phase machine and dual metric as answers to improvised multi-turn attacks, single-scalar success rates, and single-judge bias, and includes a candid limitations section on scale, coupling, and missing head-to-head baselines.

Significance. If the dual-metric gap and depth dependence hold under stronger external grounding, the work would be a useful methodological contribution to LLM safety evaluation: it makes partial versus fully actionable harm an explicit, reportable quantity rather than a buried threshold effect, and it attributes multi-turn success to phase depth via a controlled ablation rather than a single end-to-end ASR. The reproducible state-machine framing, documented technique library, dual-ASR formalization (Eqs. 1–2, score map Eq. 3), and honest upper-bound caveats are strengths relative to free-form multi-turn attackers. The phase-depth result also suggests a concrete (unevaluated) mitigation direction—conversation-depth limits—that is operationally relevant. These contributions are significant for red-teaming practice even if absolute ASR numbers remain upper bounds for rubric-aware adversaries.

major comments (4)
  1. Section 6 (Attacker–evaluator coupling) and Sections 4.3–4.4: the same frozen Combination-C jury supplies both turn-level steering (technique prune, phase advance) and the success label that defines overall and full ASR. The dual-metric gap (up to 33 pp) and the P2→P3 recovery in Table 8 are therefore success against a known rubric, not necessarily calibrated rates of actionable harm for a blind adversary. The paper states this as an upper bound, but the central interpretive claim—that the gap measures partial versus fully actionable harm—remains load-bearing and under-secured. A minimal fix is either (i) a decoupled grading scorer held out from steering, or (ii) a human audit of a substantial sample of partial vs full successes on the P4 critical items (p4_1, p4_2, p4_5), reported as agreement with the dual metric rather than only as a pre-campaign pilot.
  2. Section B.4 and Table 12: the only external grounding of the jury is a 34-item pilot (F1≈0.864, FPR≈36%) that the authors themselves treat as a sanity check, not validated calibration. Full ASR is defined by three critical checklist items (Table 10: p4_1, p4_2, p4_5). With elevated benign FPR and no reported human agreement specifically on those actionability items, the stricter gate’s meaning is unclear. Before treating full ASR as the primary metric of operational harm, the manuscript needs a larger, item-level human study on P4 outputs (including borderline partial completions), or a clear demotion of full ASR to a jury-internal quantity with human spot-audit rates.
  3. Section 5.1.1, Table 14, and Section 6: one representative goal per Moderation sub-category yields a 42-cell grid. Per-category rows in Table 14 (e.g., Indiscriminate Weapon 83.3%) and any category-axis claims are therefore single-goal signals, not category-wide vulnerability. The paper notes this, but the abstract and main results still present “seven Moderation sub-categories” as if the grid samples category risk. Either expand to multiple goals per sub-category, or reframe all category-level reporting as per-goal robustness pooled over victims and remove language that suggests category coverage.
  4. Table 8 and Section 5.4.4: the P3-cap recovers overall ASR to 97.6% largely because early-disclosure can still fire the full P4 extraction rubric under a P3 cap (Algorithm 1, dashed path in Fig. 2). The large P2→P3 jump is therefore not cleanly attributable to “exploit reframing” as a content-producing phase; it is a trigger that unlocks P4 scoring. The interpretation paragraph acknowledges this, but the ablation’s headline framing (“phase depth drives success”) should separate (a) permission to evaluate under the P4 rubric from (b) necessity of dedicated P4 extraction turns. Report full ASR and early-exit rates under each cap more prominently so readers do not read the 61.9 pp overall-ASR jump as pure phase-content necessity.
minor comments (6)
  1. Table 1 is a qualitative feature comparison; the caption already warns against cross-paper ASR comparison. Consider moving any residual implication of superiority out of the introduction so the contribution is clearly methodological rather than empirical dominance.
  2. Section 5.1.4: gemma-3-12b-it is both a victim and the Grader seat. The caveat is stated; a one-sentence sensitivity note (e.g., full ASR with Grader seat swapped) would help readers weight the gemma victim row.
  3. Equation (3) and τ=7: the score map is designed so one critical item lands exactly at the success boundary. State explicitly in the main text (not only in 5.1.5) that overall ASR is nearly synonymous with “≥1 critical item,” so the dual gap is almost entirely the all-critical vs any-critical distinction.
  4. Section C disaggregated views are labeled “ceiling case” (Run C). Flag in the main text that technique rankings (T26/T29) and phase-of-breakthrough percentages are from the highest full-ASR draw and were not verified on Runs A/B.
  5. CBRN and self-harm goals are framed as safety-evaluation queries (Table 2). Ensure the public artifact release plan in Section 7 is consistent with not shipping checklist templates that encode operational detail criteria in a way that is easily inverted into attack guidance.
  6. Minor clarity: Algorithm 1’s early-disclosure branch evaluates the current response with the full P4 rubric; a short note that this can credit success without ever entering the P4 phase state would reduce confusion when reading Table 15’s early-exit share.

Circularity Check

0 steps flagged

No circular derivation: AMT-X reports empirical multi-turn ASR and a phase-depth ablation under an openly coupled judge-in-the-loop harness, not a first-principles result forced by its inputs.

full rationale

This paper is a red-teaming systems/evaluation paper, not a derivation of a physical or mathematical law. Its headline claims (overall ASR 97.6–100% vs full ASR 66.7–78.6%; collapse to ~29–36% under P1/P2 caps and recovery at P3) are measured outcomes of running a scripted attack grid against six victims, not quantities algebraically identical to fitted parameters or self-defined premises. The dual metrics (Eqs. 1–2) and the critical-item score map (Eq. 3) are transparent definitions of how checklist votes become rates; defining overall ASR as score≥τ once one critical item passes is metric design, not a circular prediction of that rate. The phase machine, technique library, and UCB1 selection are constructive attack machinery; the depth ablation holds goals/victims fixed and varies max_phase, which is a controlled sensitivity experiment rather than a tautology. Attacker–evaluator coupling (Section 6)—the frozen Combination-C jury both steers pruning/transitions and labels success—is a real Goodhart/upper-bound validity caveat the paper itself states, and the 34-item pilot (FPR≈36%) is only a sanity check; those concerns belong under correctness/external validity, not under circularity as defined here (no Eq. X = Eq. Y by construction, no fitted parameter renamed as an independent prediction, no load-bearing self-citation uniqueness theorem). Prior multi-turn work is cited as external context (Crescendo, CoA, ActorAttack, X-Teaming, etc.), not as an author-owned uniqueness result that forces the present claims. No significant circularity found.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The central empirical claims rest on a black-box conversational threat model, hand-chosen phase budgets and transition thresholds, a frozen three-model jury and score map, and a single representative goal per Moderation sub-category. The phase machine and dual metrics are the paper’s main constructs; the technique library is largely consolidated from prior jailbreak literature rather than independently evidenced as optimal.

free parameters (5)
  • Per-phase turn budgets B0..B4 = 2,2,3,2,3 (max 15)
    Set to 2,2,3,2,3 (sum 12, hard cap 15) by calibration to concentrate turns on high-yield phases; not derived from theory.
  • Phase-transition thresholds δ_check, δ_phase, δ_disclose = 0.70 / 0.60 / 0.40
    Fixed a priori at 0.70, 0.60, 0.40 as ‘reasonable defaults’ (Section 5.1.5), not swept or cross-validated.
  • Success score threshold τ and critical-item score map = τ=7; bands [7,10]/[3,6]/1
    τ=7 and the raw-score bands in Eq. (3) are chosen so one critical item lands at the success boundary; set with a small human pilot, not independently validated.
  • Jury composition and debate rounds N = Combination C, N=2
    Combination C (gemma-3-12b-it, llama-4-maverick, gpt-oss-20b) with N=2 fixed before experiments after a 34-item pilot.
  • Attacker temperatures = 0.3 / 0.0
    Planner 0.3 / analyzer 0.0 chosen for generation diversity vs deterministic features.
axioms (5)
  • domain assumption Black-box text-only adversary with full history re-supply and no weight/system-prompt access is the relevant production threat model.
    Stated in Section 4.2; scopes out white-box and leakage threats requiring deployment secrets.
  • domain assumption Multi-role debate across distinct model families reduces single-judge positional/verbosity/self-preference bias.
    Adopted from ChatEval-style multi-agent judging (Section 4.4.1) ‘by design rather than demonstrate the reduction empirically.’
  • ad hoc to paper P4 critical checklist items (complete structure, real values, operational steps) are necessary and sufficient for ‘full’ actionable harm.
    Defines full_ASR (Section 4.4.2–4.4.3); gates the primary metric without external validation beyond a small pilot.
  • ad hoc to paper One representative goal per Moderation sub-category adequately samples category risk for headline ASR.
    Grid design in Section 5.1.1; authors note this as a scale limitation in Section 6.
  • standard math UCB1 multi-armed bandit over techniques is an appropriate selection rule under scarce turn budgets.
    Algorithm 2 cites Auer et al. UCB1 with logarithmic regret; used as a standard bandit tool, not proved optimal for jailbreaks.
invented entities (3)
  • AMT-X P0–P4 phase state machine with semantic-signal transitions no independent evidence
    purpose: Make multi-turn attacks reproducible and ablatable by phase depth rather than free-form improvisation.
    Primary methodological construct; distilled from anonymized red-team engagements, not independently measured outside this pipeline.
  • Dual overall_ASR / full_ASR metrics with critical-item gate no independent evidence
    purpose: Separate partially actionable successes from complete operational detail.
    Defined in Eqs. (1)–(2) and the score map; the gap is the paper’s main empirical product, not an external standard.
  • 31-technique phased attack library with UCB1 selection no independent evidence
    purpose: Supply phase-bound tactics and adaptive selection under budget.
    Consolidates patterns from prior jailbreak literature (Table 11); library composition is design choice, not independently optimized.

pith-pipeline@v1.1.0-grok45 · 30454 in / 4128 out tokens · 41564 ms · 2026-07-14T06:39:04.840462+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation." pith.science (2026). https://pith.science/paper/Z27HHRVP

@misc{pith2026260711151,
  author       = {Pith},
  title        = {Pith review of: AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z27HHRVP}},
  note         = {Machine review of arXiv:2607.11151}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete operational detail. We propose AMT-X (Adaptive Multi-Turn Exploitation), a phase-structured multi-turn red-teaming framework. Unlike prior multi-turn attacks that rely on ad hoc escalation or free-form per-goal plans, AMT-X casts the attack as an explicit, reproducible multi-phase state machine driven by semantic signals from the victim, and replaces single-judge scoring with a multi-role jury whose phase-conditioned checklists gate success on actionable harm. Across six frontier victim models (queried under their default safety alignment, without added moderation layers) and seven Moderation sub-categories, AMT-X attains overall attack success rates of 97.6-100% under a lenient score threshold, but 66.7-78.6% under a stricter gate requiring complete, real, and operational detail: a gap of up to 33 percentage points between partially and fully actionable harm.

Figures

Figures reproduced from arXiv: 2607.11151 by Alex Leung, Kentaroh Toyoda, Yi Ting Shen.

Figure 1
Figure 1. Figure 1: AMT-X system architecture. Every adaptive component runs inside the attacker (dashed box), which reaches the [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The canonical P0→P4 phase schedule. Solid arrows are normal advances; dashed arrows are the early-disclosure shortcut that jumps straight to extraction (P4). P4 (gold) is the only phase that counts toward ASR. Per-phase budgets 𝐵𝑖 and transition thresholds are reported in Section 5.1.5 [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: An illustration of technique selection with a [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Per-victim-model overall and full ASR (Run C). [PITH_FULL_IMAGE:figures/full_fig_p017_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Overall and full ASR vs. maximum attacker phase. [PITH_FULL_IMAGE:figures/full_fig_p019_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Overall ASR by victim across phase caps. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Two-phase pipeline: pre-campaign judge selection (top, human-grounded) feeds a frozen jury into live AMT-X attack [PITH_FULL_IMAGE:figures/full_fig_p026_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Scorer call graph from request to calibrated scalar. [PITH_FULL_IMAGE:figures/full_fig_p027_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Run C disaggregated resistance views: (a) per-(technique, victim) blocking rates and (b) threat sub-category [PITH_FULL_IMAGE:figures/full_fig_p029_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Phase-breakthrough attribution: (a) aggregate success counts by breakout phase and (b) the same broken down per [PITH_FULL_IMAGE:figures/full_fig_p030_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Technique-effectiveness views (Run C): (a) decisive-turn prevalence, (b) marginal per-technique ASR, and (c) the [PITH_FULL_IMAGE:figures/full_fig_p032_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Dialogue-efficiency views: (a) within-𝐾-turn success fraction and (b) average dialogue depth per victim [PITH_FULL_IMAGE:figures/full_fig_p034_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 21 linked inside Pith

  1. [1]

    Alon and M

    G. Alon and M. Kamfonas. 2023. Detecting Language Model Attacks with Perplexity. arXiv:2308.14132

  2. [2]

    P. Auer, N. Cesa-Bianchi, and P. Fischer. 2002. Finite-Time Analysis of the Multiarmed Bandit Problem.Machine Learning47, 2–3 (2002), 235–256

  3. [3]

    M. G. Azar, Z. D. Guo, B. Piot, et al. 2024. A General Theoretical Paradigm to Understand Learning from Human Feedback. InInternational Conference on Artificial Intelligence and Statistics (AISTATS)

  4. [4]

    Y. Bai, S. Kadavath, S. Kundu, et al. 2022. Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073

  5. [5]

    C.-M. Chan, W. Chen, Y. Su, et al. 2024. ChatEval: Towards Better LLM-Based Evaluators Through Multi-Agent Debate. InInternational Conference on Learning Representations (ICLR)

  6. [6]

    P. Chao, E. Debenedetti, A. Robey, et al. 2024. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track

  7. [7]

    P. Chao, A. Robey, E. Dobriban, et al. 2023. Jailbreaking Black Box Large Language Models in Twenty Queries. arXiv:2310.08419

  8. [8]

    Derczynski, E

    L. Derczynski, E. Galinkin, J. Martin, S. Majumdar, and N. Inie. 2024. garak: A Framework for Security Probing Large Language Models. arXiv:2406.11036

  9. [9]

    P. Ding, J. Kuang, D. Ma, et al. 2024. Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts Can Fool Large Language Models Easily. InProceedings of NAACL-HLT

  10. [10]

    Ethayarajh, W

    K. Ethayarajh, W. Xu, N. Muennighoff, et al. 2024. KTO: Model Alignment as Prospect Theoretic Optimization. arXiv:2402.01306

  11. [11]

    J. L. Freedman and S. C. Fraser. 1966. Compliance Without Pressure: The Foot-in-the-Door Technique.Journal of Personality and Social Psychology4, 2 (1966), 195–202

  12. [12]

    Ganguli, L

    D. Ganguli, L. Lovitt, J. Kernion, et al. 2022. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv:2209.07858

  13. [13]

    X. Guo, F. Yu, H. Zhang, et al. 2024. COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability. InInternational Conference on Machine Learning (ICML)

  14. [14]

    R. Koo, M. Lee, V. Raheja, et al. 2023. Benchmarking Cognitive Biases in Large Language Models as Evaluators. arXiv:2309.17012

  15. [15]

    N. Li, Z. Han, I. Steneker, et al. 2024. LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet. arXiv:2408.15221

  16. [16]

    Li, W.-L

    T. Li, W.-L. Chiang, E. Frick, et al. 2024. From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline. arXiv:2406.11939

  17. [17]

    X. Li, R. Wang, M. Cheng, et al . 2024. DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers. arXiv:2402.16914

  18. [18]

    X. Liu, N. Xu, M. Chen, and C. Xiao. 2024. AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. In International Conference on Learning Representations (ICLR)

  19. [19]

    Y. Liu, G. Deng, Y. Li, et al. 2023. Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study. arXiv:2305.13860

  20. [20]

    Y. Liu, X. He, M. Xiong, et al. 2024. FlipAttack: Jailbreak LLMs via Flipping. arXiv:2410.02832

  21. [21]

    G. D. Lopez Munoz, A. J. Minnich, R. Lutz, et al . 2024. PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI Systems. arXiv:2410.02828

  22. [22]

    Mazeika, L

    M. Mazeika, L. Phan, X. Yin, et al. 2024. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. InInternational Conference on Machine Learning (ICML)

  23. [23]

    Mehrotra, M

    A. Mehrotra, M. Zampetakis, P. Kassianik, et al. 2024. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. InAdvances in Neural Information Processing Systems (NeurIPS)

  24. [24]

    OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774

  25. [25]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, et al. 2022. Training Language Models to Follow Instructions with Human Feedback. InAdvances in Neural Information Processing Systems (NeurIPS)

  26. [26]

    Panickssery, S

    A. Panickssery, S. R. Bowman, and S. Feng. 2024. LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076

  27. [27]

    Pavlova, E

    M. Pavlova, E. Brinkman, K. Iyer, et al . 2024. Automated Red Teaming with GOAT: The Generative Offensive Agent Tester. arXiv:2410.01606 22

  28. [28]

    Perez, S

    E. Perez, S. Huang, F. Song, et al. 2022. Red Teaming Language Models with Language Models. InProceedings of EMNLP

  29. [29]

    X. Qi, Y. Zeng, T. Xie, et al. 2024. Fine-Tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!. In International Conference on Learning Representations (ICLR)

  30. [30]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, et al. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems (NeurIPS)

  31. [31]

    Rahman, L

    S. Rahman, L. Jiang, J. Shiffer, et al. 2025. X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents. arXiv:2504.13203

  32. [32]

    Q. Ren, H. Li, D. Liu, et al. 2024. Derail Yourself: Multi-Turn LLM Jailbreak Attack Through Self-Discovered Clues. arXiv:2410.10700

  33. [33]

    Russinovich, A

    M. Russinovich, A. Salem, and R. Eldan. 2024. Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack. arXiv:2404.01833

  34. [34]

    Do Anything Now

    X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang. 2024. “Do Anything Now”: Characterizing and Evaluating In-the-Wild Jailbreak Prompts on Large Language Models. InProceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS)

  35. [35]

    Souly, Q

    A. Souly, Q. Lu, D. Bowen, et al. 2024. A StrongREJECT for Empty Jailbreaks. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track

  36. [36]

    A. Wei, N. Haghtalab, and J. Steinhardt. 2023. Jailbroken: How Does LLM Safety Training Fail?. InAdvances in Neural Information Processing Systems (NeurIPS)

  37. [37]

    Weidinger, J

    L. Weidinger, J. Uesato, M. Rauh, et al. 2023. Sociotechnical Safety Evaluation of Generative AI Systems. arXiv:2310.11986

  38. [38]

    X. Yang, X. Tang, S. Hu, and J. Han. 2024. Chain of Attack: A Semantic-Driven Contextual Multi-Turn Attacker for Large Language Models. arXiv:2405.05610

  39. [39]

    Zheng, H

    C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang. 2024. Large Language Models Are Not Robust Multiple Choice Selectors. In International Conference on Learning Representations (ICLR)

  40. [40]

    Zheng, W.-L

    L. Zheng, W.-L. Chiang, Y. Sheng, et al . 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InAdvances in Neural Information Processing Systems (NeurIPS)

  41. [41]

    unrestricted researcher

    A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson. 2023. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv:2307.15043 A Phase Checklist Definitions Phase rubrics are implemented as structured checklist definitions wired into the scorer at harness startup (Section B sketches aggregation). Table 9 summarizes per-phase criticali...