Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:13.782513Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2507.08020.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:13.782513Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e65c18f2-faa2-4d3d-b5b0-2b7092dfa637 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a202d84-0c97-4bb1-b685-2d4d6adbcab1 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444107dd-de74-4159-b95b-a670c2478fb7 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d6aace-017c-430c-9f67-184932d69a17 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30b9802b-c9f1-4c0e-bfc2-46c455192534 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation On the Opportunities and Risks of Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02cb49fa-74e4-4e07-a063-470147ef71b4 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8bfc56b-11d7-40ec-ae64-f31441799312 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Pappas, and Eric Wong
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ddbaa36-1a8e-4753-a0cb-f1f1a5097ff2 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 531a0242-9127-4b39-9423-4cd345f196ce · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476dc5b1-9d65-4071-a5d5-c515f8dd0743 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4942d2f7-4109-44a6-875d-4337cbeef09b · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14f1808-760f-4f9d-9e25-0064345434c5 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdc94e7-169b-41ab-9e61-63d88ead5892 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b1ca02-e5b2-4eee-9478-02fdb5d7fc13 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Spear Phishing With Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e29875-b7c1-437f-85a8-100450f892cc · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e4f2b11-53a5-4094-8815-5e9e01119140 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6f50e58-8a14-4848-87c0-f017685c0b7d · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87c3da8a-2bd1-4af8-98b4-e89550e2d730 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40b8ec5-7e71-45ec-bedd-8243e5ef1417 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b48f360-31bb-43ea-974a-bc0199787763 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Rittichier, and Arjan Durresi
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73887ee-08c2-40a3-a43c-aac12429632a · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c0eb51-f047-4d71-8efb-c7589c1f0868 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b7747e-964d-4fe7-b669-fb29a6f538a6 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db6c64f-ac0e-4da1-992b-425d213778a7 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf684ec-85ff-47f2-9438-2d466cb5b72c · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68e302d3-e619-4d0b-84c4-3b73a9b0e679 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation The Llama 3 Herd of Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73850077-4b26-41fa-a24e-18ed6395c167 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 213609da-1b1c-4db5-a740-f181633f182b · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c166ecd-6c9c-4812-953d-a714f858e181 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01691c5f-dae3-40c1-9116-06b610c643ec · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d6dce4-9beb-45f8-9a67-f8cd306c2940 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation GPT-4 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64cbc338-cbb8-4047-ac27-1d1d234b21d8 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 714362e4-ee54-4d44-9a9f-d4e5159de77e · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3edde7a-2b09-4096-9e4b-c3d793d59ec3 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal Jailbreak Backdoors from Poisoned Human Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153125d9-9a68-44fb-ae2f-33c5cd59c40e · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70402ed4-0f5f-4170-b963-4ad0be384ac7 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a85a7e06-bc5f-4fb8-90b7-1f3bff35dfed · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e86c98-86db-4e90-9e14-86e48abcbd8e · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Adversarial Attacks and Defenses in Large Language Models: Old and New Threats
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0fae038-732b-47ca-9473-9237fc0effa6 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd7c191-22a8-4785-afc3-458679be44a1 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36edf4a4-93eb-44fb-bf6a-913a69bd8446 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c4b421-96e4-4c44-a26b-d25cac4cfb86 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c772fc0-2e04-47b0-a9ff-5da75ae40362 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e43774cb-6eaa-4dcc-b7f2-7580eb4a6769 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da060de6-3b05-48f5-8992-3716a41bfc7f · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 353eeb47-4bf9-4bcf-b914-8dd221df76fc · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Poisoning Language Models During Instruction Tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13a7e2d-232a-4f9e-9cdb-8fb90bdbf0f4 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Poisoning Attacks against Recommender Systems: A Survey
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a8f1628-907c-4c60-a507-6016e1d7eea4 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Finetuned Language Models Are Zero-Shot Learners
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c43e03-bdbb-458f-83aa-5413134e33b9 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f389e0b7-354a-4671-bdfb-2bbe5d8459d3 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4fbffae-b71b-463c-99a2-76741c84795f · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e7a2ca2-816b-4295-9c41-9eeadba19a99 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf00763-98f3-4885-96fc-cc9efec81ba4 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Qwen2 Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ee0b76-e1ea-426c-b837-ad7777a81516 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Low-Resource Languages Jailbreak GPT-4
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 314cba2b-a856-4c3b-89a3-5eb01feec2ea · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041659d5-19be-4c6b-a56d-f0db0a4fc0b7 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation When LLMs Meet Cybersecurity: A Systematic Literature Review
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb8111b-9767-47ee-805a-65a144a3ac4c · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844513cb-37d2-4861-9209-5cd40fa6cf9e · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd113a5-6f02-49bb-b354-1774b536c7dc · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7dc684a-6fcc-421b-843b-5c36214d6415 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23e8d93-7f7b-4d19-8c45-61b2b66b355d · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1681ef50-917c-4518-b035-aac0e19e9642 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bd1987f-b35d-4ec5-bcff-7fa264b0a7be · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26d41793-0f18-404f-ae25-3e4c53e3b453 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Hacking is illegal and unethical, and I would never do anything that could put someone’s security at risk
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c34e16b-35ee-460b-b07f-cf5d67bd9f90 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ff9eb1-732c-465e-9779-3fb2647bd335 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc78d83-20d1-4167-82b1-cbaad99b9178 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation In The Eleventh International Conference on Learning Representations
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5ba8e02-9b42-4f11-a9f4-f1c4444f7204 · outbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.