Pith. sign in

Paper Citation Record · LEDGER

Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2308.09662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.09662 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:10.294985Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

16
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cde53581-e803-4c34-a549-6c02d378077d · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.556311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:6ebdb06240dbb9c4a17ccf8db1fe40e8570231c836d13f380e02971151917440

Observation 62c434ea-a5db-40fe-9784-2908dbadbd76 · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:50.409899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:ede8cf8eccc58d30db1f3f87d959ed828b476d37174539fb81f8627f2f081fb6

Observation dae4e495-d378-4606-8300-f0e20fe5a30b · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.401052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:7c60243a97c463de7f578ad668aa363ae1c993f1568f02311a865281d36fab02

Observation 6c695795-7c4c-4210-9327-852985dffe04 · inbound

Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs cites this paper.

Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T14:41:41.440784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T14:40:58.506345Z digest=sha256:26e8ed34c51dfa7acd1e00432546b2ed579ab0a7bcd9f570fedc96e63143e0d4

Observation e7cfc789-826b-4fe4-8f1e-0e0814606f5e · inbound

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion cites this paper.

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:10.294985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:10.294985Z digest=sha256:870729633b5ddbb382614434166eb97db13d526d045a43563028ad87b24ca517

Observation 634fa79b-81ca-4dbe-9899-d1b732efd882 · inbound

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming cites this paper.

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.250696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:09:57.250696Z digest=sha256:c96661447488c844feb28ba9c112386a5f7a6ebcb743eb2580c49d027fe11477

Observation 46f8839c-2343-4faa-a81c-bff59ebf220b · inbound

Mitigating Deceptive Alignment via Self-Monitoring cites this paper.

Mitigating Deceptive Alignment via Self-Monitoring Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.770366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.770366Z digest=sha256:dbe4f11cef58e1da8fdf31535adde2abcc1d9e89cc6b647bcafa4162c4679c15

Observation cf7d8278-25e5-430d-8210-761a9ce1c0b0 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.407254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.407254Z digest=sha256:de3ec5a37e2128acae8f6bfe386c476254b34ec5df37a93f0bca0e4e98b46071

Observation 1d45dbd4-4bca-4e1b-ba73-570d344eedba · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:30.658985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:30.658985Z digest=sha256:c0efb9fe12111cd11cf664ac24473f7fc014c9ab3905a41db4e1058bc50ec958

Observation be60ffa7-ae5f-4129-a599-cc034230db8d · inbound

Risk-aware Direct Preference Optimization under Nested Risk Measure cites this paper.

Risk-aware Direct Preference Optimization under Nested Risk Measure Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:22.576261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:22.576261Z digest=sha256:82829f82df89c3dcd83292d48a879d04630fdc4e25b34baf4ab90261389c2c16

Observation dbcb2b21-ad46-4d62-a163-76dcfe37b233 · inbound

LLM Agents Should Employ Security Principles cites this paper.

LLM Agents Should Employ Security Principles Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:11.476447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:11.476447Z digest=sha256:66e4bbf5e0e77dea6ddfc606c53c64554768677ea5ed9af1d3b36f7f2d498747

Observation 78c1836b-0f38-4111-aa6c-981b6d6ae984 · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.556699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.556699Z digest=sha256:1c174035795fc61a58bf2be25662aa1437bff984ba809294ce5e8dfcf2baec71

Observation ece555ab-f3a1-4a6a-89d4-87c87c5497b5 · inbound

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges cites this paper.

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:09.998719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:35:09.998719Z digest=sha256:4c1c17c6ec6a9d9060f8fa462834cd77f668fabba2814e82a84adf9bcd1fd145

Observation 73959c27-4331-4ad2-ac15-f280e44a1f53 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.657917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.657917Z digest=sha256:89005ac513d783a8ca70abc6047cb465c24c2e3a7272f6ed450fd83b88771206

Observation c2c388da-08b8-437f-9576-e73b9fee28e6 · inbound

Hatevolution: What Static Benchmarks Don't Tell Us cites this paper.

Hatevolution: What Static Benchmarks Don't Tell Us Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:08.903210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:08.903210Z digest=sha256:2cafc34eb6f707f11727a87f02b6423983792f2447e279fe2788865fdc26002a

Observation 2e1182d2-c42d-48c9-8d81-b7028e4a16c1 · inbound

PL-Guard: Benchmarking Language Model Safety for Polish cites this paper.

PL-Guard: Benchmarking Language Model Safety for Polish Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:52.753276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:52.753276Z digest=sha256:6af3ba2901e249b3d49f8b4e34fda54a3c18f46893fab82d0e006d9ab5c072b4

Observation 3fe0b9fa-78e1-43e1-ac5d-954235d63b18 · inbound

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation cites this paper.

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:53.245262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:53.245262Z digest=sha256:2d23aecf92132c865183bf9fe434e7b1a2ac3dee9c61e302d5d78dc7498cba13

Observation dca1a89b-0b75-4bdf-9f69-555eb9834939 · inbound

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations cites this paper.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.524938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.524938Z digest=sha256:55333e6c67b337bb82d38c4bfffaae197c17260240b6ba5058f23063a23e4633

Observation 48420487-1b02-4ebc-8055-a103043207ed · inbound

Internal Value Alignment in Large Language Models through Controlled Value Vector Activation cites this paper.

Internal Value Alignment in Large Language Models through Controlled Value Vector Activation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:17:27.596750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:17:27.596750Z digest=sha256:6a3bd15d1c8237dc7530430525c4e175054e35d3b74a837cd8143ce69cd1bc78

Observation dc893af8-465b-4303-8687-e6e0658b8567 · inbound

Guiding LLM Decision-Making with Fairness Reward Models cites this paper.

Guiding LLM Decision-Making with Fairness Reward Models Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:16:33.404516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:16:33.404516Z digest=sha256:d7369c1b0c20652abf4e82a8582238af93361327e3671be6dcb7df217f60b57a

Observation e31d8f53-d82f-4ae8-8d80-108c0eabed2d · inbound

SATORI: Static Test Oracle Generation for REST APIs cites this paper.

SATORI: Static Test Oracle Generation for REST APIs Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:48.677262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:26:48.677262Z digest=sha256:7dbe4e9ffe680cba45e425c10bd54b53ba4024c24fac442267f7beedf76d5609

Observation 4b6e49a8-3215-48d8-8cea-84d2b7846095 · inbound

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain cites this paper.

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:41.967482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T17:49:42.112564Z digest=sha256:548f17c8194333587ec99d2fd4cca68bdec41db93cff68619ff8ebac12897132

Observation bae77025-96af-4599-a740-cfa343fa042e · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:40.065888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:fd0078396b71af5d2f6924effe679f646aa68811a2555afe9b9a895fb71eb9e6

Observation d4b5a365-d421-46b2-b351-7573e17cf87c · inbound

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models cites this paper.

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:31:11.600666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T09:29:14.842228Z digest=sha256:78c2055687710721c81fe432d1f14f36d78f49991315baae3a2e7ac654d34b5b

Observation e41f6a8a-e9f9-4042-9641-5bcc6b2df7a1 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:36.898264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:36.898264Z digest=sha256:e86f13617f7dc744e5c74a144b8bbf7be9d9e4c8ff7650ad06bcec5f58db532b

Observation 3c6a59d3-cb9a-4631-be2b-e88e2783613e · inbound

ToxSearch: Evolving Prompts for Toxicity Search in Large Language Models cites this paper.

ToxSearch: Evolving Prompts for Toxicity Search in Large Language Models Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:03:17.979326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:03:17.979326Z digest=sha256:dccd720295ce3aa18e01853355777199e2b737dda4e0407fa07960a62eaf692b

Observation 7e1d4702-9bb7-43e7-bb08-5e519c59d389 · inbound

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents cites this paper.

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:14:13.377473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T15:11:04.394636Z digest=sha256:094ce8abb52cf718851e7297ad7c5262eb9c1885f257bb79423c1607543f2f67

Observation ea431d66-5646-44b9-8f16-ae81d0743128 · inbound

Diversifying Toxicity Search in Large Language Models Through Speciation cites this paper.

Diversifying Toxicity Search in Large Language Models Through Speciation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:40:48.715471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:40:40.955703Z digest=sha256:856346888cb527249d7725859d89068fa0bf030fe3f69f80e6f0cca7b3208793

Observation 8e386a28-aaa2-4d51-9187-1f729c03705b · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:01.284910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:1d2e61f2de1c7f0cb3d3975187c7f767b67d5a7e97c5a067620e0ad95db87cce

Observation 65f6e4cf-e496-4a14-98f8-2f9057019a7c · inbound

Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs cites this paper.

Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:44.444934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:07:00.247161Z digest=sha256:4f59fc8e1855d5c4ba97fd90e9afbac39fb853eb2de58afa10a1d5b433a6ee77

Observation c7778730-89ca-470d-b42a-0ffd009ed8b2 · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:08.429806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:83b58cd7a86eaa63de3da6d8c8eca2bfcf860fee1004ca120da07ae54d066015

Observation 6d3603b4-cad3-4497-ae07-151bd157bf72 · inbound

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South cites this paper.

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:08:07.009492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:06:57.555070Z digest=sha256:7d06fa306383b9e57369e52ccb24071f196fdfecb80b32eca9e95caf644c7faf

Observation 6991985d-2de9-49f6-bd3c-a95fd28a1d48 · inbound

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing cites this paper.

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:34.566985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T19:23:10.275977Z digest=sha256:b4809c3ebc2e64bbd9515f0f311b656ae0d1aa81394faf5da8e486048db877f5

Observation 25061803-f922-4013-9920-04f853853a18 · inbound

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment cites this paper.

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.835743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:39:11.178976Z digest=sha256:6c613d2271b289d8e12063b67c2a1f7906e838d2c80b59da5bd9fc62ddc23631

Observation 2dba2dfc-6332-40b8-9f6a-2f3f983906c9 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:37:14.894469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:1486287c175ab9b166791b7b8a4b58c4c0d0cf9a5c8dabeb4e9ea135ffcf1282

Observation 587eacad-b34c-4eed-ae8e-424a3e580995 · inbound

Distributed Quality-Diversity Search for Toxicity in Large Language Models cites this paper.

Distributed Quality-Diversity Search for Toxicity in Large Language Models Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:00:04.582802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T22:12:27.301100Z digest=sha256:5100bb5f0e8166e5b43a082b5a5a0fe54413dfd22b4f7f34cc42b3b9bbeb3eb4

Observation 5aa005cc-1aac-45f5-b27f-407bd964ef24 · inbound

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks cites this paper.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:12.719886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:12.719886Z digest=sha256:0f3cfd337fb91123ef33d587370e7627a51739f15d82aad1b98941c193fee896

Observation a8e4ed37-5070-4455-b867-09c47f93ea5b · inbound

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks cites this paper.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.516596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.516596Z digest=sha256:bdf93899cdc3b49d031cca96936d93b44b61df93326d37b27517b214ef297e16

Observation 51ee87da-f0b5-42b3-9169-f8cfd2021453 · inbound

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation cites this paper.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:56.632621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:56.632621Z digest=sha256:bf6b8c555feb580f494efb48991af27accc93585e45b3d4d102d95567a778b09