Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:10:31.131754Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2504.13562.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:10:31.131754Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 62ae4d59-dd8c-457d-8afb-1fa5b705cfbc · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d58959d5-2197-41e6-8d01-c45298f405c0 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ede2e4a-7d9d-4021-8cc7-cc0ba662f677 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245933f6-3bad-4a58-866a-850cf9a12834 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Bender, Timnit Gebru, Angelina McMillan - Major, and Shmargaret Shmitchell
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e1e37a-e73a-4df6-92d2-fade10287c68 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b600ed47-243c-4441-adb5-d84d93d8382d · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fffb375a-25d3-45e4-a074-0cfa3666257f · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd13752-c144-4f96-9926-123a15e39aa2 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f908982-cb68-417b-9abc-8fb5f5cf95d9 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Injecting Universal Jailbreak Backdoors into LLMs in Minutes
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca86a27-66a4-4eb5-b4bb-a387c53d98ba · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90b09ef-8ac4-4e71-876e-b7d460e28410 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00dda28-cf7a-4a73-a981-9336b9192a41 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f625397d-fc9d-4e2d-9788-082c094a3b0e · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d3d7ce-ea34-46aa-b1fb-04b9522bfce0 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1389b44-281d-4d50-b791-e7350b403733 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Mixtral of Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac9a438-846e-4cff-9e05-9a54cf1bcfb3 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb49374-8d37-42f5-a66b-36c1df2b2d4d · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b46e0dc-c6ca-4f1e-9976-5d8a59afcfcd · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3493a01b-a5f8-40d1-8da2-7247ba2ac36c · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ba5215-cad9-43cf-9589-6974ce7b167b · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3965006-66f1-4af7-b424-54cc5c34482a · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9402466e-5bd3-43c8-ad65-98a5b4b713f9 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dba8af0-03c4-4526-a6d5-e5399af47a86 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cca35e4b-f552-4d59-99dc-e326b1bf4058 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09df1fb0-8d16-476d-bdac-33c129394e40 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification GPT-4 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77fca6da-6759-4b32-a6e5-4e44c4cc73df · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 91c07911-987f-4f36-af22-692348e74303 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28163107-d7c6-4216-b8a1-b3b2d3f08e8b · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22002250-cb8b-45a8-88bf-b3ecec84cc53 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce3cdde-fb4f-4623-922d-eba00cc09e96 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification do anything now
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c84cdb34-fce2-45c7-afff-22cb6dacb3af · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbfe326f-9e8d-4b9f-8e54-c1bbbfbd1e94 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c50aca7-63f7-4a47-b6ce-969c24fcdf55 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a3f5d5-f0c5-4208-a783-e54ec33b01d8 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ad9db7de-fa27-496f-bae5-2d5b87b42f9c · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c3d0bbf9-03d8-49cd-8153-bb7f999c5e68 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b454e88e-88f5-494e-bbc2-cc822f24a2ef · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f809632a-e5c1-4ace-94e5-516b24aafc35 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a701a33-ae22-4dda-9ab4-a90d4ad77df7 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fbc9df5-e981-4a39-b09a-557393605066 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa10b5e-15c8-4546-b981-9aa37f58c252 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0e1ef83-954b-4e3d-baf3-6fe8801ae300 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df04aee-21bb-46bd-9156-c2910edfe740 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a167b720-56ec-45c0-b909-80cdf83c8234 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f68c0b31-1ae0-4ea4-a11b-c12c54df7d38 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Instruction-Following Evaluation for Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8fdd32-ceb5-46cb-98db-96a416576729 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a81a26c-05cf-4e47-ab71-dfbc0f7b0720 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Don't Say No: Jailbreaking LLM by Suppressing Refusal
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e069761-60d8-472b-acb5-1e6bbb1abf55 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e724aed9-0c47-4b6e-bdbe-b1b18cf33114 · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification online" 'onlinestring :=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc4eada2-5e12-4294-ab6c-4d89b87174ba · outbound
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification write newline
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.