Pith. sign in

Paper Citation Record · LEDGER

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2507.08020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08020 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:13.782513Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved62
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e65c18f2-faa2-4d3d-b5b0-2b7092dfa637 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.240046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.240046Z digest=sha256:2b5d3113547a9698b79ce018f86f744dfed42c9c9ede649342a3060852f41a24

Observation 7a202d84-0c97-4bb1-b685-2d4d6adbcab1 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.325795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.325795Z digest=sha256:d6ae2e066af4398f7315b53b3caf4b353218973f20420d103c88ee6ab16232b6

Observation 444107dd-de74-4159-b95b-a670c2478fb7 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.448232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.448232Z digest=sha256:7694a5980fc1fb66680775da4ca464d39406c9265eb768d9125ff884e9946259

Observation 88d6aace-017c-430c-9f67-184932d69a17 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.767093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.560615Z digest=sha256:f326ce5e31ad3294cb92749a8c43bf6f6f764ff14b79637537d4075fa8e38c91

Observation 30b9802b-c9f1-4c0e-bfc2-46c455192534 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.617300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.617300Z digest=sha256:f477dfbadbbd73b757c572c1843a7e20c9b2acb54ce968375b5625b9e1d0eecc

Observation 02cb49fa-74e4-4e07-a063-470147ef71b4 · outbound

This paper cites Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.620368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.620368Z digest=sha256:bd9726f19ec63e30bd0a0c9da7a31de24330d623f73abc050761dcc2c20fba32

Observation c8bfc56b-11d7-40ec-ae64-f31441799312 · outbound

This paper cites Pappas, and Eric Wong.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Pappas, and Eric Wong

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:26:14.759006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.623324Z digest=sha256:39acd3c1596906f5ed3cdcbfa44614792664a9588b8f883cccbacf0a188c605b

Observation 6ddbaa36-1a8e-4753-a0cb-f1f1a5097ff2 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.750722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.628809Z digest=sha256:53b26e199440b42a7bc408459a647235b4085565019fe7981a9f9717896c3eb4

Observation 531a0242-9127-4b39-9423-4cd345f196ce · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.631030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.631030Z digest=sha256:90430e079b880227f4da3e831da4d482bf5ca7a55bd30094c646675e805445e2

Observation 476dc5b1-9d65-4071-a5d5-c515f8dd0743 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.633838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.633838Z digest=sha256:bb01441c8b2ee4d26e90c387a6e3e400c44b64c6256cfce0aa82e7d5a6d5a4d2

Observation 4942d2f7-4109-44a6-875d-4337cbeef09b · outbound

This paper cites Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.636474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.636474Z digest=sha256:ca60add5ba6bcbd099e26dc4822d01a61e5931420b44a68619d31382e5da1467

Observation e14f1808-760f-4f9d-9e25-0064345434c5 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.639279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.639279Z digest=sha256:82dc45f258df6cea10724d283b6a2476f8a9311db0174921e632a99f1184ca69

Observation 7cdc94e7-169b-41ab-9e61-63d88ead5892 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.641819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.641819Z digest=sha256:fa34a1b767d706492972527bb3ff8ddd30eba5455f1c2bf8e7454f677ac94d6c

Observation 80b1ca02-e5b2-4eee-9478-02fdb5d7fc13 · outbound

This paper cites Spear Phishing With Large Language Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Spear Phishing With Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.644775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.644775Z digest=sha256:78ee13b13cbbb025348233c44e7e6f6512a681f0ccd76a3693aa56bbc9e33ba6

Observation 98e29875-b7c1-437f-85a8-100450f892cc · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.647771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.647771Z digest=sha256:d202ad1bfee3ba446fa475fd5cb6d0ef794e8b75fe4ef03425a04c67743d86e2

Observation 9e4f2b11-53a5-4094-8815-5e9e01119140 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.737212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.650243Z digest=sha256:7ea2444e07512c11ff8d649b2191c88fbc3d5a904b6bec5bdb87c67e0011fbb3

Observation c6f50e58-8a14-4848-87c0-f017685c0b7d · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.729005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.655580Z digest=sha256:1b33c367918585eb75881591deec530f34262b3c8f19694aee387957fc333067

Observation 87c3da8a-2bd1-4af8-98b4-e89550e2d730 · outbound

This paper cites Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.657815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.657815Z digest=sha256:64fd8e123bcb9cf509cc3fc5f71eeaecd4ef6ac2c899965368fb46d3891e8f32

Observation b40b8ec5-7e71-45ec-bedd-8243e5ef1417 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.660236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.660236Z digest=sha256:cb3e71c38008c5508e5d2116918b14e7e79ffe51e107acd2d5bf963314d21b6a

Observation 3b48f360-31bb-43ea-974a-bc0199787763 · outbound

This paper cites Rittichier, and Arjan Durresi.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Rittichier, and Arjan Durresi

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.662534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.662534Z digest=sha256:02094e4a10d2604f0cefb339f699f61a707dcd7dba338b98320e6a5f6d09ebf4

Observation e73887ee-08c2-40a3-a43c-aac12429632a · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.665283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.665283Z digest=sha256:3ecd38d7ceca02756c7dd08868d4fc1ea45b45457708c2c69f2649c03d8daa70

Observation 94c0eb51-f047-4d71-8efb-c7589c1f0868 · outbound

This paper cites Model-Editing-Based Jailbreak against Safety-aligned Large Language Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Model-Editing-Based Jailbreak against Safety-aligned Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.667591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.667591Z digest=sha256:fcee15cf4756cdcdff007844c3ae5d7d8bd1ccd6b29d5cc2acd9dc70900db5c6

Observation 42b7747e-964d-4fe7-b669-fb29a6f538a6 · outbound

This paper cites Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.670577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.670577Z digest=sha256:9d1776644ba6fbbc4daebd8be652568f83bac673635bb074b53f8892076c4ed8

Observation 6db6c64f-ac0e-4da1-992b-425d213778a7 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.673300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.673300Z digest=sha256:84a6d6c2478a072a3056fe600f45dc6b7b3c2c3f7708a6580667e7c71c6cce67

Observation bdf684ec-85ff-47f2-9438-2d466cb5b72c · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T19:26:14.715425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.676210Z digest=sha256:af18e0d02b8d16d30a57d036254ead812e8162228651956850f26d68ea8a8e5b

Observation 68e302d3-e619-4d0b-84c4-3b73a9b0e679 · outbound

This paper cites The Llama 3 Herd of Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.678386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.678386Z digest=sha256:03f4dc4c976051123658d9d17130224c25fe16139379e471ffb6d6cca2aa1594

Observation 73850077-4b26-41fa-a24e-18ed6395c167 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.707705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.680832Z digest=sha256:89be55ceac293ef097774e008e06a291657a24754bfd7b76549a7efe1a202d6f

Observation 213609da-1b1c-4db5-a740-f181633f182b · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.699821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.683041Z digest=sha256:99be56c6bc10f1511056cfabfa81f1bbbb7a537e3205c8c30476e5d699e22602

Observation 2c166ecd-6c9c-4812-953d-a714f858e181 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.691654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.685528Z digest=sha256:86aa1ab24b2b53eb6affffba0d7c1b31fa0da61fbccaa21bfb395ac16e692b85

Observation 01691c5f-dae3-40c1-9116-06b610c643ec · outbound

This paper cites Decoding Secret Memorization in Code LLMs Through Token-Level Characterization.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Decoding Secret Memorization in Code LLMs Through Token-Level Characterization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.688117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.688117Z digest=sha256:69751781ef14ef058be3472c538eeac4211873f26b6a576e0dc3a79af63ac813

Observation e3d6dce4-9beb-45f8-9a67-f8cd306c2940 · outbound

This paper cites GPT-4 Technical Report.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.691187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.691187Z digest=sha256:006878505fa7c07e6332d56528c4b75f189339b86bff36c6c3ac3ade671a5880

Observation 64cbc338-cbb8-4047-ac27-1d1d234b21d8 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.684118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.693897Z digest=sha256:5a9837db3d6db6f31be32df821905733d28f0c72631445a34cb74582b8065033

Observation 714362e4-ee54-4d44-9a9f-d4e5159de77e · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.676617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.696362Z digest=sha256:bd2ee0db5c7190aa29d1c6d6a260d8b704cb6183557388e3d8339a4804d0f98f

Observation a3edde7a-2b09-4096-9e4b-c3d793d59ec3 · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.702267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.702267Z digest=sha256:b7433578189bb781133e0f277add8404a1cf3791282f8b10c6d3ff77cb451c8f

Observation 153125d9-9a68-44fb-ae2f-33c5cd59c40e · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.705071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.705071Z digest=sha256:8bf54deea09390c40f14d11152270dff59fceb3b54f48cb040fe94207da6a112

Observation 70402ed4-0f5f-4170-b963-4ad0be384ac7 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.660638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.707736Z digest=sha256:fcabeee26da5fd1beff0b2f8de3ac09718753bf6aea3a2bd3e3a01ea6b36348a

Observation a85a7e06-bc5f-4fb8-90b7-1f3bff35dfed · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.710725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.710725Z digest=sha256:fe23a4be5efbbb06121ce84876b9ac123239a641f010afa5158fda8807875c14

Observation 56e86c98-86db-4e90-9e14-86e48abcbd8e · outbound

This paper cites Adversarial Attacks and Defenses in Large Language Models: Old and New Threats.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Adversarial Attacks and Defenses in Large Language Models: Old and New Threats

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.716377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.716377Z digest=sha256:1ba1ed03714987c9f738957f5ba6125a7c91dfde113ebf013e3e2f061392f25e

Observation c0fae038-732b-47ca-9473-9237fc0effa6 · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.718850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.718850Z digest=sha256:784b7c18eda2616094e400280953edf9dd906b0f9c721701b7cbb0bb0b20e0f6

Observation cfd7c191-22a8-4785-afc3-458679be44a1 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.647814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.721307Z digest=sha256:d6ef696dc2560678056ea7f239fd99144baed1bc63147d1fa71bd6a476c41ac2

Observation 36edf4a4-93eb-44fb-bf6a-913a69bd8446 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.723589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.723589Z digest=sha256:799e573b287330107a9d74dc24386634cdcaae20c04cd1a58ae32beff3cddf90

Observation b1c4b421-96e4-4c44-a26b-d25cac4cfb86 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.635425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.725753Z digest=sha256:7f491d700d2339e4cfedc5b9f3c67df51b93a95944dc48eb5336f768a56fb8ba

Observation 7c772fc0-2e04-47b0-a9ff-5da75ae40362 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.627959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.728119Z digest=sha256:297e1f79f9ca6390f370b8cdd13f7787aeec81dc0e1bc5647612f857cb4394e0

Observation e43774cb-6eaa-4dcc-b7f2-7580eb4a6769 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.730532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.730532Z digest=sha256:f37580a8342d0cf0a52b4bd1a7b240752c4d0fe5911e12e91c2d96d75dfc01e2

Observation da060de6-3b05-48f5-8992-3716a41bfc7f · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.733803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.733803Z digest=sha256:4f7649e2ebb7b924f3fc3e0c80f4bfe7627725e8df2edbc63db62cf220604a15

Observation 353eeb47-4bf9-4bcf-b914-8dd221df76fc · outbound

This paper cites Poisoning Language Models During Instruction Tuning.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Poisoning Language Models During Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.736202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.736202Z digest=sha256:294e400930d076bae59a0ab53ef4be39445c9c3eee274f1568dc989d81f6e106

Observation e13a7e2d-232a-4f9e-9cdb-8fb90bdbf0f4 · outbound

This paper cites Poisoning Attacks against Recommender Systems: A Survey.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Poisoning Attacks against Recommender Systems: A Survey

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:26:14.151515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.738590Z digest=sha256:b60f4500a6fd30cff69bdc8a8cb313b1af318763e3fcd87260f891d611130a22

Observation 0a8f1628-907c-4c60-a507-6016e1d7eea4 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Finetuned Language Models Are Zero-Shot Learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.741057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.741057Z digest=sha256:0105017b4bd2507324b1d806d759208c3cc3508485f9f5b60627e97cc7e073ad

Observation 78c43e03-bdbb-458f-83aa-5413134e33b9 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.743717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.743717Z digest=sha256:78fc6ddef506266ec3d845cdadc5627cdd663b30410286e3ef7a7ee9f61f8144

Observation f389e0b7-354a-4671-bdfb-2bbe5d8459d3 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.746263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.746263Z digest=sha256:911064556a25d86952dbb8d772ff80c006bae7b46766d3ea8ae2d1d341930343

Observation a4fbffae-b71b-463c-99a2-76741c84795f · outbound

This paper cites Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:26:14.125247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.748695Z digest=sha256:cecd0fee51087919fa911ce3783540d95ccd6710fd9c09155593d82188dea7d8

Observation 0e7a2ca2-816b-4295-9c41-9eeadba19a99 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.751199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.751199Z digest=sha256:b741aa7a23fe4918b975e2e5f4686a46ff54d7bf733145e5c1decf81e380a7f1

Observation adf00763-98f3-4885-96fc-cc9efec81ba4 · outbound

This paper cites Qwen2 Technical Report.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Qwen2 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.753557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.753557Z digest=sha256:8c7d4903c3a370b68b9f9307687704affd8ee44b2a48555aa37fa565af2fe54b

Observation f7ee0b76-e1ea-426c-b837-ad7777a81516 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Low-Resource Languages Jailbreak GPT-4

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.755908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.755908Z digest=sha256:16f8300f5097e0eefc0d31230fb4c9283c9ce6b3973822e3e2dc6e40b2a652ce

Observation 314cba2b-a856-4c3b-89a3-5eb01feec2ea · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.758637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.758637Z digest=sha256:d056d87c40067f6ab4372568cedb7bdea56fec7e899771b6659f847c891a0233

Observation 041659d5-19be-4c6b-a56d-f0db0a4fc0b7 · outbound

This paper cites When LLMs Meet Cybersecurity: A Systematic Literature Review.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation When LLMs Meet Cybersecurity: A Systematic Literature Review

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.760965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.760965Z digest=sha256:2d0c792d6930bd9f478e6214f9b26a62495f8cced02fa7d99005b2df56793056

Observation adb8111b-9767-47ee-805a-65a144a3ac4c · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.763688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.763688Z digest=sha256:449c4dc33c9c143e31e3c3ae900851e7ae0f324aa27eb2d3dd56c90d75d6611d

Observation 844513cb-37d2-4861-9209-5cd40fa6cf9e · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.766148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.766148Z digest=sha256:518903bf9dfbb76db37ac60615cc2db943c67501209d4ddec341206747ef4024

Observation 7dd113a5-6f02-49bb-b354-1774b536c7dc · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.768941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.768941Z digest=sha256:fb52ea13cfebdf560dbd28e354a324d56783852397aaf7137c935068c9cc29de

Observation c7dc684a-6fcc-421b-843b-5c36214d6415 · outbound

This paper cites Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.771379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.771379Z digest=sha256:ae705cfc1e17f4452536d23383f4ec5e2ef7d0c5c9856174f24c6ff228ec4f79

Observation a23e8d93-7f7b-4d19-8c45-61b2b66b355d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.774858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.774858Z digest=sha256:cf84b4fc7e86282e4935a9725f72ac31a3de7a454dd446ba8586b8d64dfff77b

Observation 1681ef50-917c-4518-b035-aac0e19e9642 · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.610511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.777653Z digest=sha256:05551b8016df481d9e023d29f7c83a92b8867caddc007ea1ad0ce446e0db9dd0

Observation 5bd1987f-b35d-4ec5-bcff-7fa264b0a7be · outbound

This paper cites an unresolved cited work.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:26:14.601969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.780234Z digest=sha256:6fbc19c5a5e41717b4868a24906f18f50bcc17d6eb6f82805c5a33bf9cef052d

Observation 26d41793-0f18-404f-ae25-3e4c53e3b453 · outbound

This paper cites Hacking is illegal and unethical, and I would never do anything that could put someone’s security at risk.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Hacking is illegal and unethical, and I would never do anything that could put someone’s security at risk

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:26:14.593579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.782513Z digest=sha256:09f84a393b7f656fbd9cbac9ae996a96bfdc3818ffaf8aec96c5d6f0b12659ce

Observation 8c34e16b-35ee-460b-b07f-cf5d67bd9f90 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.713457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.713457Z digest=sha256:56002c13f13b7d26e37c28ed8ea489008ebefb9eb1b962782d34e3fa174e5725

Observation 61ff9eb1-732c-465e-9779-3fb2647bd335 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.373964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.373964Z digest=sha256:15b3eb7d7e39848b42b79efc02db0cb1150f11e055404c0470ae344b3d880db5

Observation afc78d83-20d1-4167-82b1-cbaad99b9178 · outbound

This paper cites In The Eleventh International Conference on Learning Representations.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation In The Eleventh International Conference on Learning Representations

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:26:14.668413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:26:13.699124Z digest=sha256:32f149946aa861ec34195e4b52ac9a43de3f7aeaf5446d83bf9462c3bd88087c

Observation f5ba8e02-9b42-4f11-a9f4-f1c4444f7204 · outbound

This paper cites Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.652637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.652637Z digest=sha256:72dbc3d813d96b057fb38e078b4f9ea62cd275475f777a5a872e4da7a5be5149

Pith citing papers

No inbound Pith citation observations are available.