Pith. sign in

Paper Citation Record · LEDGER

LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2408.15221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15221 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:11:36.094864Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:30:00.600665Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00f3009b-b943-4e6c-a81e-c9876dd4efe8 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.196729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:657725a6a82b1f5ee0ec8be56c168dcfdc282aab62572891e768f066c526d954

Observation 6be63de9-ad04-4610-934b-efa8f3402205 · inbound

Eliciting Language Model Behaviors with Investigator Agents cites this paper.

Eliciting Language Model Behaviors with Investigator Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T16:11:36.094864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:11:36.094864Z digest=sha256:8876c3ac1e87df976a1f3117ecf85ecaf99374dd4e1330d017b5f56284e90448

Observation e6cdc523-6da2-4ed6-b067-4bd784cf81c9 · inbound

Position: Adversarial ML for LLMs Is Not Making Any Progress cites this paper.

Position: Adversarial ML for LLMs Is Not Making Any Progress LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T12:47:21.695561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:47:21.695561Z digest=sha256:78bb3e49bd2e128929f0a13aee0edfe7efb0af91eec58dadeb72e623c4fe9474

Observation e85b16c9-7c66-4b34-93ef-ddfa98ce1114 · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.176057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.176057Z digest=sha256:32d5843dd4021c86c3e77a06363dd9b4525ff58c0367c721c8f230e773558d3f

Observation e498973f-982f-4427-82d7-4087f05a47f0 · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.290970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.290970Z digest=sha256:a4bfd18d1b9022ad3721a96b6365ad625ab91c6e0ef1c813441729e228d9b733

Observation c9d9c45c-9194-4078-a496-571e02d37ab6 · inbound

Fast Proxies for LLM Robustness Evaluation cites this paper.

Fast Proxies for LLM Robustness Evaluation LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:32:24.045500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:32:24.045500Z digest=sha256:20f9d5b7ad846eeea8f587917467c12941c210a6a48e9ace1ba8d7ed4d4c68f8

Observation d3962040-2111-4a15-87eb-fd55a83b7fa1 · inbound

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming cites this paper.

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:58.868771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:09:58.868771Z digest=sha256:a535377184189a968cdb555696b59eae2cf660d86afdfe68844ced44cf60d968

Observation 025da3cb-08c6-465b-bc04-ee7fc410a902 · inbound

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects cites this paper.

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:35.050363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:35.050363Z digest=sha256:5f81fad5195b3e286fbc0cbfc3c03176242a642c18de7de29ac24d9920fdab55

Observation af5baf9c-24f2-4c7e-bfbf-4d60fce44742 · inbound

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues cites this paper.

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:46.808275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:46.808275Z digest=sha256:7ad5a7a4ad10ec976e4b7867ee9dd34fa2e06b9e069c8d4eb1771b74c79ba541

Observation 0795bf9e-e8f0-4fc9-b33f-0846d3705edb · inbound

Existing Large Language Model Unlearning Evaluations Are Inconclusive cites this paper.

Existing Large Language Model Unlearning Evaluations Are Inconclusive LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.420996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.420996Z digest=sha256:2d438eb6a5daaafc22b400193d50c7b36e70280a3c9adfc1cc057237819a1bc3

Observation ea5b7d3f-4fcd-41a9-bb98-f17d2047398d · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.490378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.490378Z digest=sha256:548a9cc53b7cc70eb36ca872a7d77885434362a622456d1ad5c9a46c0972f23e

Observation 423fc428-ce25-4c67-86de-47204dbbddfe · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:20.823652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:20.823652Z digest=sha256:c7a6e80e1ef52ef99393084e8dd33d971b4fa4cf4600a86b25cde7f78c55db63

Observation d5622b38-d1cb-47b7-9f96-0a23a7be94db · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:01.464128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:01.464128Z digest=sha256:b7d2d76580c1aa6e08e2ef65d9d8468e47d45395c3e6f05045495b506130ca2f

Observation c6e2f740-a3b4-4596-b545-b38fd8c10bda · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:24.381051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:24.381051Z digest=sha256:a5c1cb2a55c7a4d15cc263dc3f2814e14796c70f4dcb9dd0973b3418f76cc499

Observation dc7aef9d-0703-43cf-9d52-107eb2ead49c · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:52.389653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:52.389653Z digest=sha256:7ce85a3457124c362d18beb3bf282ccec25109386ac6d36435f1135d82f6b994

Observation 0f2e6602-08ca-43d2-9b71-c8070ef34c97 · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.687570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.687570Z digest=sha256:6f0ebbdf4fc5ef700d7aec1d5d1f7a87833cee6fc26daa45b8960154f5167bbc

Observation 309ec9ad-75ad-489f-8785-ee5c42109a05 · inbound

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks cites this paper.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:45.154452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:45.154452Z digest=sha256:83d3d9da7d871b311b8e20bb6d8b6c7594c5fb5cc456d484d33b5289bf5f2623

Observation 4b3e7a82-fdde-47cc-9155-3044d9651a7f · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:46.079982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:46.079982Z digest=sha256:ebde89adcc8aa0f3075bd789e8b46f255201c90bba99476c75f06acc9f9634d6

Observation 1a126705-4d3b-41e5-91f1-f56f557d7299 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:55:38.102301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:ed5d12c35df9a3166699f43c648b7e2c8f8eec34b5d6fe4db1fee9f957d557ee

Observation 247a5f25-22ba-4dda-8c19-878517708f42 · inbound

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense cites this paper.

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.926434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:26:19.922383Z digest=sha256:50ae21677b3fd241ba14e95980d5734c656c9b0012a81e083b19abe705eabf2e

Observation a0f878cb-8114-447b-940e-cd4f89f64fbb · inbound

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion cites this paper.

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:04.298032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T15:19:46.920899Z digest=sha256:77f1e678cb6e8b675c6e21773020049d17abadf9f6d2633ac9defd2e1d72094a

Observation 7536afa1-8f93-42f7-aebb-3233d15663a0 · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.319458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:d43e946e71eee106bed3b6f5cab1b72a665b3c6c1c92f0bebb3603dd629dc9e2

Observation b8f758b8-04bd-4c98-b15a-60c8b6524bd3 · inbound

MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks cites this paper.

MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.234082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T05:31:47.682478Z digest=sha256:b38085a783053830dd11fda5ba43db1535a0f497a67ce1bc08f42f165facf287

Observation 564491a8-4c05-4997-9103-3bee6e9e20a3 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:01.117969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:84ba668a057bad0b55a21f6faeec4898c307120f9c844eeb3edbaaf081505d00

Observation dfc78e53-93fb-41db-bf6c-48a286a0781e · inbound

Evolving and Detecting Multi-Turn Deception using Geometric Signatures cites this paper.

Evolving and Detecting Multi-Turn Deception using Geometric Signatures LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:23:32.700901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T15:17:58.804973Z digest=sha256:978810b9dd7ac6456f200f21fb9bd58d53de42d34b3d8da3a4bd5de94fa0387a

Observation 9946a1cd-f8e2-4ad4-82e5-f7b21607fdff · inbound

Learning from Mistakes: Can LLM Self-Recover after Misalignment? cites this paper.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T18:51:10.298187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:51:10.298187Z digest=sha256:1e895d9f3860f8b50a49100fdcd6d1979d963e31bc43c96515ce7f8c9cefd49c

Observation 284104f1-8e15-4221-b954-329445444012 · inbound

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models cites this paper.

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:56:20.699869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:51:17.667213Z digest=sha256:cf206d70cd9d44a1d5479287845dd23634a521ef141768ab2c45b71643425b52

Observation 3049dbee-2d34-46df-a22d-a70daf20d4b8 · inbound

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents cites this paper.

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:16:36.009627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T09:15:57.044886Z digest=sha256:ffa7b155cb050329b0c85a7490a1b2e6de6a0fa48a2f7affad46427407af99a9

Observation a3a43767-df8f-4547-8278-1b7e140699c0 · inbound

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models cites this paper.

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:18.134705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:51:34.829877Z digest=sha256:79d54a99ca9c30948440690a6e3358922af264f42c626482dae25e24e58c23dd

Observation 61326c0e-fcbf-4c6e-9818-5e9ed641c4ea · inbound

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails cites this paper.

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.534360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:42:05.886984Z digest=sha256:a0e2a5d7e9bf41a90c0a6b6cf972bba660c9528e36dcd90cf213b42fe7ecd9b2

Observation 69b52be5-acb9-44ee-891c-ecefc9fbcba6 · inbound

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms cites this paper.

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:34.476904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:13:26.201462Z digest=sha256:380cda4c64133a37ea218be011939bf7249e4f7d92602943fa184da006ded935

Observation 621dde8d-66f7-41e2-8abd-fef7d8c646a7 · inbound

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models cites this paper.

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.528094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:52:02.327522Z digest=sha256:1519686f136593a24abd3de311c37c052894200835e8e209635eabac98ed2b59

Observation 61d50880-a3b8-44fd-8e0d-baf2b1480e20 · inbound

Do Thinking Tokens Help with Safety? cites this paper.

Do Thinking Tokens Help with Safety? LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:30:00.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T23:37:49.412578Z digest=sha256:c497c2f0cbbd5e1723398c128b9a3d31dd8d9970c13e308583188cbc15eac0b3

Observation c78087e5-adb2-4f36-8087-e8dc34bc18db · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:54.925819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:4058ff5aa03e79c7fe5f22495ecce34a41a5d408cf69c57f7f15346ba19fdc14

Observation 368139f9-4c32-473b-8a35-3f186dd78d3a · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:d9332953129ab7dc71d0b33990d15984cf16fb6d0749fc95e69c3f00d0b53aa7

Observation 52be3610-9d22-4093-b472-d455bb546b5d · inbound

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security cites this paper.

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 1939

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:08.626130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:19:08.626130Z digest=sha256:7aa548cf0e9524497189f1842c420e93c51ed64d643ab38d64840c366695dd88

Observation 2b05152e-9391-4fb2-8bf2-2754d6943dd3 · inbound

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems cites this paper.

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.458741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:54:34.458741Z digest=sha256:48d869ff87a0c2b2b42a3fdc23c8c925678bd256ad5c49cdd947c6f668193393

Observation 517c58ae-e33a-4b57-a802-dda793e6ad6d · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.309791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:24.309791Z digest=sha256:5eccd39425e1199b3b69c988685c83387d5c367dace0a05430535558e695aa4a

Observation 0067c403-d465-4008-9400-5843a1aea210 · inbound

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion cites this paper.

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T11:50:47.128838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:50:47.128838Z digest=sha256:b12dfe63194b2077f032de927682965e99d654327c574361de8ffb86bcd97830

Observation 51f6832f-3f61-4b2a-8830-0224e9affcb8 · inbound

SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks cites this paper.

SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:00.935942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:32:00.935942Z digest=sha256:15fa148318644998a1f1d080bf4a37cd86c640ad18c3e537f406c3ee810b3bda

Observation fb46f4d2-7cbb-4a41-ab8f-5d17ec9f892c · inbound

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation cites this paper.

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T05:44:22.263222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:44:22.263222Z digest=sha256:91bb11794a5b84e1db1d86390271029fd8f1e3aa10fffbb0a3e99ef43a4cfc50