Pith. sign in

Paper Citation Record · LEDGER

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

As of 18 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.20766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20766 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:12.986775Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2aa4da68-6b4f-4008-b654-dcbe48a3917c · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Yi: Open Foundation Models by 01.AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.761630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.761630Z digest=sha256:6c586aa647d310e5f7510fdc15c66eab10459f9e7b4045a837b0620a2715890e

Observation 092afef3-3837-4b5e-8f95-9b1ae6b685bb · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.766265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.766265Z digest=sha256:04ef7b9e1741ff787d5876b8b175307434505b48a4164dee16de062f1ba059fc

Observation ff2285ec-1da2-4be1-9a5b-cdc9c5650f98 · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.770432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.770432Z digest=sha256:48329dc5955a77635cdabbc7320ebc9b9e9ba2465e2e973fedf1fef537fa0a40

Observation 468cad06-0e4a-414a-8c22-166bb21232c4 · outbound

This paper cites Towards inference-time category-wise safety steering for large language models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards inference-time category-wise safety steering for large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.899160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.774158Z digest=sha256:a7633dc5da697a01ec522f5260faf8150a7b8427e57ae5ca5feb6d79421cb21c

Observation 8333e408-0fca-4c09-8e88-c446a4ecf603 · outbound

This paper cites Man is to computer programmer as woman is to homemaker? Debiasing word embeddings.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Man is to computer programmer as woman is to homemaker? Debiasing word embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.777705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.777705Z digest=sha256:26ff6c8afd14c91b68275e4e5c563244616627decaea74555c21ec14d665f93b

Observation 1e117786-9786-498d-b6c0-fdacfa42fa86 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards monosemanticity: Decomposing language models with dictionary learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.781141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.781141Z digest=sha256:ca19e2cb55f4a1cb1b652f49728e11522e637c3c14fe3641d4406ab53dd4c9e1

Observation 35e38398-c650-48ed-bb6e-9c72e74f13cd · outbound

This paper cites Language Models are Few-Shot Learners.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.784667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.784667Z digest=sha256:52ecfd80c73be1f5ebe6bee984fda37094acc9a85b9e00c351791f65178d9905

Observation 17474615-b21a-4117-bbc4-ac63ae9ab4fa · outbound

This paper cites On the Measure of Intelligence.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection On the Measure of Intelligence

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.788366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.788366Z digest=sha256:7ad41675b05d95610eff4f89503d3085a4252fe45520e80ca4826b6e5cc160bb

Observation 4cd21b06-c7ea-41e7-b016-ad9b2b1c5695 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.791829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.791829Z digest=sha256:29eb8b017fa00fc4669356accf61868e5d1f995ccea097c0a537bf5f90792120

Observation 66b00bbb-6d0c-4ff4-b25f-da59a5d282f7 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.795167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.795167Z digest=sha256:eaca71ea2eff5fca0ec8e70d4e611d3fa8f15a24340c00e8213148911b614d0f

Observation 51f2370b-d7be-4c1f-aeff-02c1b5730df9 · outbound

This paper cites https://dphn.ai, 2025.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection https://dphn.ai, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.872091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.798641Z digest=sha256:5f9e3aa742b53836ae4c81a0435c636ab18e69f456278f00139ee7ab4d56373d

Observation 2c782017-4040-435c-952c-0013deccf568 · outbound

This paper cites Toy models of superposition.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Toy models of superposition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.801847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.801847Z digest=sha256:37c234cfe83d9e1c284bf4d91a98357eda26125b025fb29c2224fb03e9ff265f

Observation a12d6ed4-a3ec-41f5-b817-e37bfd8aa8aa · outbound

This paper cites Finding alignments between interpretable causal variables and distributed neural representations.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Finding alignments between interpretable causal variables and distributed neural representations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.856147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.804855Z digest=sha256:36b95c1e87525a5be9e2c1d3851818b9d3e02dc51b1942e20c162bb64856c236

Observation ae0c54c9-a499-4401-81f9-d061c0c906b9 · outbound

This paper cites SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:57:13.715920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.808005Z digest=sha256:83352fe19b206f57dfc8844160af06190598585220b7e1ec9960fe4770525599

Observation a2f0ac79-d72b-4a3a-b0fe-910bf950ab7e · outbound

This paper cites A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.811315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.811315Z digest=sha256:596f87c50729ef623768f21ed2edc53e0d92dfb9939c7818924590845e53256c

Observation 4271ee28-0264-4c95-8ded-ef9226004dbb · outbound

This paper cites Model merging and safety alignment: One bad model spoils the bunch.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Model merging and safety alignment: One bad model spoils the bunch

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.814839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.814839Z digest=sha256:fe92226f6ffbd2e0d7fd2a8cfa8375345de7b054c2799042a8f6cfae568032db

Observation 47b3b0e0-4190-4f7b-812b-b3ae8e6a6549 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.818090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.818090Z digest=sha256:da83bea011d26e22b59ee4e20e976a29dd2e723ed946771789728cb2b1602a68

Observation aad500bd-6164-4250-bb38-159e43dab79b · outbound

This paper cites Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.821622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.821622Z digest=sha256:7af07cc5206cfb29575a9e30855ea4cadcea7b35dd981d43d961b06e1b0f6a36

Observation 33e8ee80-4730-41b1-b93b-9511a0285653 · outbound

This paper cites SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.825477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.825477Z digest=sha256:d913e1b79ad0a27ea2b7a94fb7b94f46e785453aa1590040002a7b324ca04be0

Observation ca596ab6-2639-4d8b-8f7c-4b3ed707e016 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.829123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.829123Z digest=sha256:bf9988d63133627189be4eb932910cf0f35769d9c54b0e3a534be7782871edc7

Observation 1d963b07-a33a-4bee-b4dd-929f49162067 · outbound

This paper cites The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.832694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.832694Z digest=sha256:cef166a4dc76dc768435360df378e6c1bfbbac6a95fc0566b4372226568bb02c

Observation a161f0dc-f6f2-4fcf-8c5e-44d44cfe8f82 · outbound

This paper cites Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.836202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.836202Z digest=sha256:e7484781e643375354b8486e3753f25ed758f4bd0cc50ac04657ab54afcf5be5

Observation 302948ed-b1c0-45e3-b42f-698298bbc876 · outbound

This paper cites What makes and breaks safety fine-tuning? a mechanistic study.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection What makes and breaks safety fine-tuning? a mechanistic study

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.845319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.839957Z digest=sha256:46a87039c1d829b7f7fdb0468c34a403c80fdb2b9b0a16c0e7f2e1cec1ce4f94

Observation 99bad9e1-902a-4446-b2a1-a415beaa3217 · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.843060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.843060Z digest=sha256:4ede61a248c46148d7c35fe5216f2bd594bb62a1fa4b922bc569faec28b9822c

Observation 0fc4d333-17c4-4fbc-a229-329185933c1b · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.846665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.846665Z digest=sha256:5e793706866a1a9b28a1754e52aef6ad54df28930af6d6ccd19aed4e29238916

Observation 755422c7-39cb-4e43-a9dd-d8e7670c5380 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.850194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.850194Z digest=sha256:6a5b4db8394f1c90030cdb0620ed69d79a68fdc0d72cf3db52c853a332cec9c1

Observation 6d9662f8-177f-4f6a-9e85-d35648a8fbd0 · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Inference-time intervention: Eliciting truthful answers from a language model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.854241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.854241Z digest=sha256:b909775a8695b7591d007b84c736c305d27ef9a259c92f02e0c10264795a2e98

Observation cfb2e00b-e697-449a-a39c-6d85ae6cf5ae · outbound

This paper cites Rethinking jailbreaking through the lens of representation engineering, 2024 b.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Rethinking jailbreaking through the lens of representation engineering, 2024 b

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.857443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.857443Z digest=sha256:b94431eca1ae5a62c975bb793b21b8f0320f49e7efb57630d8309daa027679ff

Observation 1d3844fb-c2cb-4f88-91fb-61e375962b98 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.860810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.860810Z digest=sha256:fd3204dae4b01c8931883db2eb6f54b789361f0dee5b17f10e588bf61cf6c770

Observation 6023b781-e395-45a4-8076-f3037ddade19 · outbound

This paper cites Towards understanding jailbreak attacks in LLM s: A representation space analysis.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards understanding jailbreak attacks in LLM s: A representation space analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.864421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.864421Z digest=sha256:c07d9631c9596161392706f6f1005066fa4f55152f7db356ec5a956482e03d56

Observation d28e1915-567e-4f34-87db-0b7f74f29fb1 · outbound

This paper cites The Llama 3 Herd of Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.867614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.867614Z digest=sha256:687daf99d8daf02a8b0fd59e46db8ed89edc90bc1c4d487877a7c6f274c6a1c4

Observation 486f20dd-9218-40a3-abae-e940bb414b97 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.871465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.871465Z digest=sha256:08f11fdf5b958e29adb062c7a5ed5d1cb38b9bc40458549ea42f0b4becce8247

Observation a6f9c9c6-d705-4f4e-baa3-cb5f4dfe2c17 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.875001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.875001Z digest=sha256:8fc5543b7943c77767f9f621da95afcd0cec2fe4f656cce4518387e7301ccde2

Observation bb034246-6af4-4a59-b1e4-b960e64aae42 · outbound

This paper cites Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.878381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.878381Z digest=sha256:6b0ebe761fe4fc05d57ac04d954ab753bd45e4b5acde4b46d944a4b376899f48

Observation cf2db5bc-92c8-4a3e-aba5-1c89c5bf2c6d · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Model Refusal with Sparse Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.882025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.882025Z digest=sha256:48b26eb53cb7e7853807a49e02f247cfaa4f95489bb89e13c7fb67b1b0a78ec9

Observation 486399a8-beeb-467e-bcff-766c87154cd1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.885402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.885402Z digest=sha256:838f35e3f0c05925cfa0d523300564c51f02c08b765c198e5286cf67af3025a8

Observation acd42cd7-bdda-4659-9abe-4c29ecee0864 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Llama 2 via Contrastive Activation Addition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.889009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.889009Z digest=sha256:6b24f60f8a5d338296ee0af5d2cd26db2d52db53015aa1cdc970b3f7759861bf

Observation 57a135cf-df8c-4683-85cc-9bd623cf73d4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.892755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.892755Z digest=sha256:8ee2d868eaa21730bf1e47d540042372f7c2bebd6d92b68757d602ab6c5171f5

Observation 0671c2ed-033c-458b-b1c9-eae1ae50a15c · outbound

This paper cites Qwen2.5 Technical Report.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Qwen2.5 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.896258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.896258Z digest=sha256:6f8348bfcb39ac5ba4e7ecf4dee14c998af10d98da9e83b0c526eb2f4896bf76

Observation 2ea3a346-6727-4f0e-a0ac-fe9f4cf288e2 · outbound

This paper cites Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.900075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.900075Z digest=sha256:c53c72adbd88c545dbcf3244511a8a6ac5fc3d1330402c387b2af3c43c3a0933

Observation fec3c9fe-56bb-47cb-8083-252dbc476f9b · outbound

This paper cites An embarrassingly simple defense against llm abliteration attacks.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection An embarrassingly simple defense against llm abliteration attacks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.903476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.903476Z digest=sha256:9fec450b764bac6aaaaf540d990058647b528a33624131e19b9941e7d7f97a92

Observation 97459c1a-bf2f-4334-9e57-cff5bdd29dd2 · outbound

This paper cites Hashimoto.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Hashimoto

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.906779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.906779Z digest=sha256:5b56e9f3c183f90dd5ff997438fc596fac6e2a86ada699446a887fa2ab704e21

Observation d92c241b-2aba-4afe-b9b1-ae0890ac495e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Gemma: Open Models Based on Gemini Research and Technology

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.911134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.911134Z digest=sha256:ef71a7641a7609d4552c0a6e411bc63dc368fd0831c56240cc90355123ec8b9b

Observation bdda36b2-8d12-4090-9c2c-91c91c5d6d0d · outbound

This paper cites Daniel Freeman, Theodore R.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Daniel Freeman, Theodore R

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.915277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.915277Z digest=sha256:e8c96898c534481356a8af6cd899493ae38422ef6508bb134b325f2f5ccbd0ac

Observation 99dd246e-57ef-45d8-b199-ff87eda4e38a · outbound

This paper cites CodeJudge: Evaluating Code Generation with Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection CodeJudge: Evaluating Code Generation with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.918674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.918674Z digest=sha256:ef760e390a18fd3c377ecef9f84a4f94b3d7ed96245b12469324c755c1a7d7f7

Observation 8b876ede-e860-47df-a75a-7f728f26ebd3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.922406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.922406Z digest=sha256:f8b2adecfdc6ff542fe8c2cd9815c834a07c3a36853201968cea16f3f2694c8d

Observation bab913e3-618f-44bc-8a36-de6a9893441d · outbound

This paper cites Steering Language Models With Activation Engineering.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Models With Activation Engineering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.926010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.926010Z digest=sha256:a617606cfec227f950791cd8e0c7c0114150191971ef51d29cfe51eea71f77f9

Observation fdcaea98-406f-4a1e-8a4a-c9766de5394c · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.929314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.929314Z digest=sha256:8ff7a634bd7eb07a3e6d7ae2d57a1a0cdcdfa4bd4b267c7c84f5aee54ac0e50a

Observation 6f9949d7-4e83-4504-be8a-a821d8024bfc · outbound

This paper cites Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:57:13.810919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T14:57:12.932659Z digest=sha256:e576de4d308ff5956b83f2daa814da9ccacee786d745847d4a91a9a423e44df9

Observation 853dddd6-f99c-4b48-a081-8870c95abc88 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Jailbroken: How Does LLM Safety Training Fail?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.935639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.935639Z digest=sha256:29247e4ee01bee0343e94ce13805ce19c709ad9dfcf1e7cd9cfec4198f07dd58

Observation 83bd386f-8001-44d2-b4cf-dc27d23048cc · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.939155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.939155Z digest=sha256:72c0839cfa0d3c4fd08037b1fe658df1a911b462d646dbb083d56b8ab49b72ac

Observation b9ab188a-2c4f-45f4-b452-5e3f2a36c52b · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.942505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.942505Z digest=sha256:eb4a2e2959c0987de1fef9c5b736fddeab72885ef8990fb4b8a8746acd2a462e

Observation 55d8626b-c077-46ae-b78c-1ecdbeb34004 · outbound

This paper cites A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.945857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.945857Z digest=sha256:55ff5f68da5b454ea599e1acb736a4755adc3c2f816135425182e1eeac0a65fa

Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.949136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.949136Z digest=sha256:a85a68c7b24223bd4e4ad73581c9994943927f6db630f9a4a346682f5a8b2050

Observation 2e8c9d3c-aa7e-4eb7-9ed2-fbbabcde399c · outbound

This paper cites Representation bending for large language model safety.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Representation bending for large language model safety

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.952574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.952574Z digest=sha256:964039589ff0305771924a9b0a60c5d270bfc0ebf5121d87a6e8fb2280b427dd

Observation 50435e64-1521-4101-838f-da39347e0c24 · outbound

This paper cites Robust LLM safeguarding via refusal feature adversarial training.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Robust LLM safeguarding via refusal feature adversarial training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.955790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.955790Z digest=sha256:52eccccf09e8e0d12c5ef18faaf5c0a354a87b08609f9f945747e98ad11f827b

Observation 6d65efbe-17ed-44bd-b2c3-53e3c656be66 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.959186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.959186Z digest=sha256:058fb53f11dc91998146a64ee36218349e4deee2072eb459ad39ec6a5cda9c7a

Observation cee53ca1-667e-4e37-bb0f-bb69aa9df89f · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.962460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.962460Z digest=sha256:690c8fa49e83e3046eda4e6db5da83af9ac5581a99a9be769ddffb515c38befe

Observation d37d7c29-99b8-41d1-a143-1a3e66ed316f · outbound

This paper cites Adasteer: Your aligned llm is inherently an adaptive jailbreak defender.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Adasteer: Your aligned llm is inherently an adaptive jailbreak defender

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.965956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.965956Z digest=sha256:6d805d540554cc79cb68e2b07c32df3390680dbbbb4456b283579b5aca570262

Observation c96e3b12-54fb-4287-967a-e210244f832d · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection On Prompt-Driven Safeguarding for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.969212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.969212Z digest=sha256:2858c149cfbb273b01a1b92937728fb4d15505cbc5d6314898280dd18aa280c1

Observation bf27ccbb-f785-4237-bce1-f84d60fdb599 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Representation Engineering: A Top-Down Approach to AI Transparency

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.972601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.972601Z digest=sha256:40ee0a120fee2a8413e20344905b28290873a532e8deb560ee49dbad555c95f5

Observation 2082d8e4-ab00-4930-86be-38c92635280c · outbound

This paper cites write newline.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.976053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.976053Z digest=sha256:f2ee6f034c6c8b8e296df9ed9870986551c699139d58e02dc2b990cfa5150c0b

Observation 9341250d-b03c-4094-a3f2-8a4ebbbe4621 · outbound

This paper cites @esa (Ref.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.979904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.979904Z digest=sha256:de1e6444b58c24b7ef0d9d4e6f374296ef565b9f3a17c2f7661cc4c9e7fdb363

Observation 47659d76-c0b0-4841-aa01-f10e23eebb01 · outbound

This paper cites an unresolved cited work.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.983359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.983359Z digest=sha256:e397167c6366f01eb61747054bfba8890adde784c1e07b904c8be45c141159b9

Observation a1294e0a-e785-4981-bb46-198da192b190 · outbound

This paper cites an unresolved cited work.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.986775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.986775Z digest=sha256:f7ec3894e226d2237d61c7dcbd4a1a56882b3a852baf23597b3089df66b83dd4

Pith citing papers

No inbound Pith citation observations are available.