Pith. sign in

Paper Citation Record · LEDGER

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks

As of 14 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.00134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00134 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:16:15.489997Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 798a21b2-7217-4fa7-b796-04582796ca4b · outbound

This paper cites GPT-4 Technical Report.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.264836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.264836Z digest=sha256:a6417508f0ed4822e44a78843c1ab089f173c3da1eebbc161dbe749b845f5ab9

Observation 7e9400a6-d005-41c6-add6-9971e971b577 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.358040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.358040Z digest=sha256:b87841c32f6ee8d9ad8444b4c5485ce284e9eba431b1c3e1ebd1e17f3bc9f53d

Observation 197cda93-1f6e-451b-ba6b-589f4480435f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.397390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.397390Z digest=sha256:9ac3d28465c92c07798459876952cf505a52e0be43e0fb71e5deb13b2de9dbb8

Observation 1afd53bc-d838-4a72-8bd4-2cb3df258454 · outbound

This paper cites A Survey of Large Language Models in Medicine: Progress, Application, and Challenge.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks A Survey of Large Language Models in Medicine: Progress, Application, and Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.467147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.467147Z digest=sha256:d1749bfa5bcfeded21fd43f3fce659161154dab802a94c534280fe0926de9801

Observation 155f7a77-09a2-4a6b-b4de-b421612b25a3 · outbound

This paper cites Exploring Recommendation Capabilities of GPT-4V(ision): A Preliminary Case Study.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Exploring Recommendation Capabilities of GPT-4V(ision): A Preliminary Case Study

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.515819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.515819Z digest=sha256:c6171ecab875452de06d50f734a28f814784272f85990f3b18abddba0e7653c3

Observation d1e15201-e534-458f-aad3-c61eaa37143d · outbound

This paper cites {LLM-Fuzzer}: Scaling assessment of large language model jailbreaks,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks {LLM-Fuzzer}: Scaling assessment of large language model jailbreaks,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.563310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.563310Z digest=sha256:8f8bfd3ad5e07186abd03f6e0685d81cfdcae3adc9044752dc6ddcc59159c8fb

Observation bb0cadd4-0fc1-4fa6-ba87-396222896a5e · outbound

This paper cites Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.626615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.626615Z digest=sha256:3e9ebf76f6ca534962b428bb2a2a2c2a711baaf709c9ad56cdd32dcf4bd12129

Observation ab6b3cef-02ed-44c9-aa4b-6c404eb1e212 · outbound

This paper cites Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.673549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.673549Z digest=sha256:87162a79b22cb3f429eabc49f0fa698fdbb3da6f257256fa962ee208d42b34ea

Observation 986c3a0b-3af2-4f0f-9beb-c9e6bdff0a39 · outbound

This paper cites Helping big language models protect themselves: An enhanced filtering and summarization system,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Helping big language models protect themselves: An enhanced filtering and summarization system,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.743694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.743694Z digest=sha256:84c46215f0da92c61c886cb285c74b1d87d62251cd4ddcdd3821569b744d2f61

Observation 404c6bc3-e694-4d71-8157-9785c1c560ff · outbound

This paper cites X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.808925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.808925Z digest=sha256:4b1fb897daaf3d36c6c67c55fe40f8dbdec0ddf1e8819f35cf31b1fd97ce36fc

Observation 42419e43-111b-46b5-987d-646f2b11d314 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:09.952167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:09.952167Z digest=sha256:fa3de2d50ea84c269447b4f84bf66eb3314b87fb87d0e645171b1cad9d0489b9

Observation 37abdc89-91e0-4d96-b872-5b08253372c3 · outbound

This paper cites Datasentinel: A game-theoretic detection of prompt injection attacks,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Datasentinel: A game-theoretic detection of prompt injection attacks,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.059041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.059041Z digest=sha256:d9545f9582055ba7a5fedcc6833900277345b8c13eee4a9e49634f7fc507678f

Observation d9fddcb5-449b-45cc-b026-56190ae5a93b · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.099322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.099322Z digest=sha256:7b9b3d7c478c56f048f72303f9e5756b95155c66e63e75adbfa7bdce243a8097

Observation 61e51d5b-99a8-4d04-bdca-02bb060bf374 · outbound

This paper cites Fight back against jailbreaking via prompt adversarial tuning,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Fight back against jailbreaking via prompt adversarial tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.155570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.155570Z digest=sha256:bddf2c0ec67a7eb224d6cdf8251cd8bc374f77c2bff96caa3def55b4e393fafa

Observation 6b501988-70d3-46de-a021-4fa32eacd20a · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.204874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.204874Z digest=sha256:1d7a199cea0f0d7199ff37584e338e6ff7fd84acef0e0278f1e992be81694cb2

Observation 7a4e43e2-0e5d-4de7-bdcb-23599c9c7c44 · outbound

This paper cites Distributional preference learning: Understanding and accounting for hidden context in rlhf,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Distributional preference learning: Understanding and accounting for hidden context in rlhf,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.255483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.255483Z digest=sha256:037182356dc7e62f84f2c7325d07c41dcaadd3bd0d643d623cc8fcfa4318a4cb

Observation c5232ba9-5422-459d-ba88-940486fce466 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Attacking Large Language Models with Projected Gradient Descent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.330416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.330416Z digest=sha256:7836d1ceb082bafe08c995df174640a6821f7a7d561e90ef86e256fc7a9db5b2

Observation fa11cee1-adf0-4aaf-9489-059349dd9982 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.406423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.406423Z digest=sha256:5c8d248b87773978a1d7c3af1568e86a949e0219c7ffef1bda4690c83dde38d8

Observation 0f7201d3-4551-4f8f-a646-d595f553792f · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.469528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.469528Z digest=sha256:30ef264d1858f5d865cf2724a91271b7673cee221a0c372b0610af5e16dcfdaf

Observation a36c6875-97c0-47a9-8bd8-25e0b0f97ecd · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.569808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.569808Z digest=sha256:1939c0e5849602dcdd00d9ed9f2087033432fa4cc4a7da424f6108ef3afc7e62

Observation a37eead2-fe61-4568-bb8d-e8cebb9cacac · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.686916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.686916Z digest=sha256:5bcfa636b54aee3ccb3933c466694a2fc5e6255d9993058157dda6fbe1397ce0

Observation 15e36a56-ffb3-45ed-9880-b7a8ba38d7b5 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.758446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.758446Z digest=sha256:e330f2cbe94eb96bce3192200df44d25808e0a2caaa2b87873537c7973bc4d5b

Observation b286444e-f098-4bb3-affc-415bf51b6227 · outbound

This paper cites Honeytrap: Deceiving large language model attackers to honeypot traps with resilient multi-agent defense,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Honeytrap: Deceiving large language model attackers to honeypot traps with resilient multi-agent defense,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.811489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.811489Z digest=sha256:8ec6a94219575c5f2f26ea5d511cc27f4c5c42f3f2db99324396fa5b1e0c385d

Observation 75924873-2eeb-426e-8578-a3e07ec5c661 · outbound

This paper cites Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.875255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.875255Z digest=sha256:581a0df437d46e229934f9ab28b7cfc35ef52920e4ff737718a22945c0d17de3

Observation f2c38328-d87e-4d06-bd39-e9b7c027bb6d · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Tree of attacks: Jailbreaking black-box llms automatically,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.989189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.989189Z digest=sha256:97c6da2d4d43adcc6164efdea1bfc71ed33356b6d703fd36b9561f70e3dcfe68

Observation 698e5121-2960-4aa4-a4cb-2b58e6633457 · outbound

This paper cites PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.117001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.117001Z digest=sha256:7231f7a29a353120a71fe76d131aa6e713d4a3f03521dbe1efe417d893eec773

Observation 6b57bb15-a2bf-48a0-aead-4011b5f951a5 · outbound

This paper cites ” do anything now.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks ” do anything now

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.197104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.197104Z digest=sha256:7a3dcc3ebfb775e03f42d38c05e94dd5a6351e28142d22b43619eacb75d9c56f

Observation 1e0a7775-d09c-445a-b87c-59792a2412ed · outbound

This paper cites Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.264145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.264145Z digest=sha256:7eed575458a217d6160ef0c0e31b073698e0b303357e1c67b499369b065149ad

Observation ed4c2d9f-cf68-4431-816e-f8271dc0c47a · outbound

This paper cites A Survey on Trustworthy LLM Agents: Threats and Countermeasures.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks A Survey on Trustworthy LLM Agents: Threats and Countermeasures

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.444009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.444009Z digest=sha256:837fff6d51730cb9b6991625a8132830e3364288734f39342ab970d28fbd8ac5

Observation b701a1c7-75e3-470a-8697-8a2126d352c8 · outbound

This paper cites Masterkey: Automated jailbreaking of large language model chatbots,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Masterkey: Automated jailbreaking of large language model chatbots,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.540024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.540024Z digest=sha256:e5d71bff515406e181b250b7f6883e9365f11b3912581b2ff1af706c605522f7

Observation 9043d682-2db1-45fc-bc51-2d569acad0a4 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large lan- guage and vision-language models,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large lan- guage and vision-language models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.664743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.664743Z digest=sha256:cd1be2ace534e752aa2ec0a6887341c31acb285ba55f057624d15f35b293a5e8

Observation 1292fb00-17c5-4c13-b6d4-8fd53a9db6bc · outbound

This paper cites Gpt- 4 is too smart to be safe: Stealthy chat with llms via cipher,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Gpt- 4 is too smart to be safe: Stealthy chat with llms via cipher,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.766046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.766046Z digest=sha256:1d4b48b1fa4d915c92bfe5e440fa64ff8fba880c8c2875f6c5384375864b527a

Observation e08f1a45-122a-409c-b3b1-99c7d60ad914 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:11.894746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:11.894746Z digest=sha256:cfbfedb729138d98fb32ab26c28d27a6f6dccb9bf9af43774b53d8b5b9f65818

Observation 43d83782-7851-4cf1-97a9-6032f993231a · outbound

This paper cites Boosting Jailbreak Transferability for Large Language Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Boosting Jailbreak Transferability for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.004751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.004751Z digest=sha256:11c58e79cb38981df7eea69110f914d300c6d82768405434be5906d26f2fd282

Observation be428ec5-f14b-4543-a2cc-a2b6aab04e02 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.144831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.144831Z digest=sha256:5729e63d66884fc30cacedaf141f749e3b079a5c8beece40a92f1e3771c0fcd4

Observation 9481f41b-60c7-4599-9b37-10358127434f · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.245800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.245800Z digest=sha256:0724b726473999ea711bf6dadb8bde92d6cd53d252a557b9a6ebe961e682dac2

Observation bec4a2cb-5209-4cbe-9ebe-e9c849fb47d7 · outbound

This paper cites Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.344757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.344757Z digest=sha256:00f501d2b41985d1059e327f6b3e0286a92a2b3a41f33e4fb062f644348c22b4

Observation 5b40f7fa-4ed7-4e8a-ad95-f43854b50ad1 · outbound

This paper cites Defending LLMs against Jailbreaking Attacks via Backtranslation.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Defending LLMs against Jailbreaking Attacks via Backtranslation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.515641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.515641Z digest=sha256:dafb9582e3ba448ae6ab329f21ed851d568c59ccc494eb150222cd1448a2fa5c

Observation d520fa31-9bdf-4ace-9186-f1ddc6454d72 · outbound

This paper cites Jailbroken: How does llm safety training fail?.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Jailbroken: How does llm safety training fail?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.672284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.672284Z digest=sha256:c14a5ba022d141440732a1feb4a119426e6411138f5d0be6604d2e7689a471b4

Observation 5aa005cc-1aac-45f5-b27f-407bd964ef24 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.719886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.719886Z digest=sha256:b0092feaeb7ad22a64b7edddbc1efa98724fc8139909d50bcee1232a3372b044

Observation 7a38e1ad-70e8-40d5-be3d-d16c2fb0d05d · outbound

This paper cites Attack prompt generation for red teaming and defending large language mod- els,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Attack prompt generation for red teaming and defending large language mod- els,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.788834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.788834Z digest=sha256:b86faad9374d07c352ab6db47fcba5723a8d4c45189c9dec22a66a975c0ccd9e

Observation 48601139-73e1-473a-ad66-da94405e14d2 · outbound

This paper cites From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.873272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.873272Z digest=sha256:c00e7a16e6888fbdbe56ed5fb729f318956f1f9395a5297602b7d10e40f67354

Observation fc595a05-2b83-4101-a1c1-46c4445fa5de · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.949365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.949365Z digest=sha256:c42040d7dc24b88c31df14c3f59982a9c5521df9aa031c3cc89717d20357317e

Observation 771bd9f5-e888-44a8-a10d-77ae5df37a46 · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Certifying LLM Safety against Adversarial Prompting

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.014747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.014747Z digest=sha256:f18e2f5d7f44f4c78795842a22a91eeb7d5df47706b7188037bde7561d789490

Observation 81cb2df2-f37f-457d-a8d0-16e043624665 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.134915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.134915Z digest=sha256:a6df203aacf92c64f7e747fd81c199367aa827bf91ceddb46c8fd13199aed56a

Observation ba8886ca-b654-4e61-8022-fc96553d7efc · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.252157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.252157Z digest=sha256:b9fdf2d06318d3fe9020ae951831d08d66d7c2212d059f7ca63e89dd2d982fd0

Observation 9744ff32-f663-4470-a466-7f4a778c6c67 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Defending chatgpt against jailbreak attack via self-reminders,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.364750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.364750Z digest=sha256:c0e8614c1c1174ca05c45ad72319e601aa7bb82d2e2a3963813df38858d834ef

Observation b0ef8a60-ca68-4930-a283-77ed33a62d93 · outbound

This paper cites Steering dialogue dynamics for robustness against multi-turn jailbreaking attacks,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Steering dialogue dynamics for robustness against multi-turn jailbreaking attacks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.535734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.535734Z digest=sha256:10c645e00d418e3eb5ab3f844ad8ad2829328f032beda1e201b13c945b0c303a

Observation 9921ae4a-ca99-4243-9263-0365000a2123 · outbound

This paper cites X-boundary: Establishing exact safety boundary to shield llms from multi-turn jailbreaks without compromising usability,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks X-boundary: Establishing exact safety boundary to shield llms from multi-turn jailbreaks without compromising usability,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.651328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.651328Z digest=sha256:6726fe9e12cb15a5341c78a25adaaf3787d6239eb6fa4c7082690b0ced083ede

Observation b15ac3db-aece-42e1-aa4a-759eb576099d · outbound

This paper cites RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.794829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.794829Z digest=sha256:eef916fde718b42dfbaca08eb7d392e86b30252317693de4e50fd151e4150ae7

Observation 1e43c751-4a75-44ae-8c4a-12950bd73985 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Generative agents: Interactive simulacra of human behavior,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.872894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.872894Z digest=sha256:9022d0af9bfa7879e6ddbc82d761cd5bd61a56622ca4fc184d91cfc95359c92d

Observation 6bb1203a-fa28-45f2-88ae-bd7312b242a7 · outbound

This paper cites Training Socially Aligned Language Models on Simulated Social Interactions.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Training Socially Aligned Language Models on Simulated Social Interactions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.946073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.946073Z digest=sha256:c253e8e0ab37618823a4c3883f42ddf218de0b6126bd6c5fcd6fdd4bae9f3fe3

Observation ac613f5c-476e-4f49-936a-dd1cf9cfb4f0 · outbound

This paper cites Camel: Communicative agents for.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Camel: Communicative agents for

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.022585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.022585Z digest=sha256:a93c5fc4d9778a5cef8532d98e9041fec8b879944b7a03c0843a6397e2329b22

Observation 8a872541-14d7-4d58-8867-cd96d55ade56 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.080163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.080163Z digest=sha256:0313e8f7cb1f229c37f25f9fc4168ff8eb42a2974330b2ac0db29e03cfc79288

Observation d3c7f28f-25fa-4bd5-9330-81387c2e42df · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Metagpt: Meta programming for a multi-agent collaborative framework,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.172690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.172690Z digest=sha256:9aa893fac37d126d425629798afaf709ac7ed4d3d124335933bdf5c0c2dd474f

Observation 66ee22d6-3d85-4913-9552-7bfac0e5507c · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks ChatDev: Communicative Agents for Software Development

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.275406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.275406Z digest=sha256:1753bd88d38c5ed1e6c075df228d858198f431ba4ba93c66d9ac367ba7516661

Observation 8892c0c9-d9e9-4521-b8d7-f7b5f9749de4 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.424749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.424749Z digest=sha256:361b1ffafc4e6ab1fa70ed5eea192dd01013f50a70350ee5e771686b8ccf79b9

Observation c713df20-a4d7-48cd-9945-5ffdffa5331d · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.553943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.553943Z digest=sha256:5fd6a20a7b8420aeab64237d5ca42ce28b089c35742bf4271abfe145a3c5da30

Observation ccbcfc1e-1960-4ffa-af48-ee77962be44e · outbound

This paper cites Defending large language models against jailbreaking attacks through goal priori- tization,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Defending large language models against jailbreaking attacks through goal priori- tization,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.624835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.624835Z digest=sha256:821581e6f2f62ae4789b934a42c98a431d5e7f00e5c0dc116b0cbed4e4a203e2

Observation 3de10f62-1386-4ca4-9a38-e8c914f54690 · outbound

This paper cites Robust prompt optimization for defend- ing language models against jailbreaking attacks,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Robust prompt optimization for defend- ing language models against jailbreaking attacks,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.814826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.814826Z digest=sha256:63eb5cd51df2f0306078e3546e1585b920d2618911e2593c80c6b865e058b47a

Observation e49c6208-7c87-4509-8609-7c19ad7b1c60 · outbound

This paper cites SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:14.920942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:14.920942Z digest=sha256:d92f9faaffc29489277ccb9649fab10b110affa9a010101a40b88b05c5f93d87

Observation ad2bd6db-372d-4528-bcb9-99a440857d4b · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Fine-tuning aligned language models compromises safety, even when users do not intend to!

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:15.024609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:15.024609Z digest=sha256:de3ac30f174abf56202375c9f2871c8ff28341a24ff87c38a1e0c49c2abfe903

Observation cfe5471f-fd88-46ed-a81b-d8486bf54dc1 · outbound

This paper cites Jailbreaking black box large language models in twenty queries,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Jailbreaking black box large language models in twenty queries,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:15.148273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:15.148273Z digest=sha256:ceaa7b9eb246ea92258c0e00365556c715a88ea6f9fe70c29e1ac3181350f7c6

Observation d2829aa8-a185-4f8d-bd83-8e7ce496b17e · outbound

This paper cites M2s: Multi-turn to single-turn jailbreak in red teaming for llms,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks M2s: Multi-turn to single-turn jailbreak in red teaming for llms,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:15.272601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:15.272601Z digest=sha256:2bfdcbb47fe6842bf977e7f4b9722d9b1f90f1731cf0fb529611925884ce71a3

Observation 0a22fc85-a5fe-456f-b43a-4d89dc2add11 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:15.396819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:15.396819Z digest=sha256:8113474894f9965e3ad5ed530e4b6bca4ca37160a09bf67ec3bcad5b09a7e984

Observation 1702fe08-30d1-48ab-94b9-b7ac696391da · outbound

This paper cites Cosafe: Evaluating large language model safety in multi-turn dialogue coreference,.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Cosafe: Evaluating large language model safety in multi-turn dialogue coreference,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:15.489997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:15.489997Z digest=sha256:1acebe6ec4d1c3625793838fadc54170e753007eb4f5481b716b6ab4acd66d09

Pith citing papers

No inbound Pith citation observations are available.