Pith. sign in

Paper Citation Record · LEDGER

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2310.02949.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02949 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.377658Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.659817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52fff216-cd20-4ab6-bdac-3a7b0be0fd59 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T04:18:27.647836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:42b82131e75bb986fdc45ae88932344166cdf8959e5416e54c0f9f8310b53336

Observation d7a14b8b-4a05-474c-aec7-8cd21659b5b8 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 202

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:47:56.163266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:68a8131a0c79eae474d6c1baff1893d72967485e8b7c225bd30bee24a287c71c

Observation b20987fe-9db7-4022-a645-8d8002f9c10b · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.444563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:fc514ccf0dfd185426461c063d3b7bfc2c07bcb36ed41bebefe6b5b1e813ae53

Observation 1107c248-3eca-47d0-ada5-8b11aaa71381 · inbound

Learning to Ask: When LLM Agents Meet Unclear Instruction cites this paper.

Learning to Ask: When LLM Agents Meet Unclear Instruction Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:13:28.074522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T21:08:42.276002Z digest=sha256:f2479b70c31f6b11db9b70b9f8f88aac8a14458b670db1907fa3be3c5d9d4f04

Observation e7154fa9-a96d-40ee-9d84-d5124b36211b · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.176221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:baca7e1ef3527da8dd4b87266828b10399ff00b5007d2620c6fc42d519a8a432

Observation 81b9a09f-61db-463b-8b6d-d0322382fd37 · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.132795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.132795Z digest=sha256:8b894f8f6081348d085135ef8e62bc67be27524406abc871525e68e569be08c6

Observation da6e0ecc-79bb-47cc-80ab-863b8edc129c · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.377658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.377658Z digest=sha256:38d94b9c1bceaa1bf27f060c16064127158f34fc0bc3d541921006960412274d

Observation 2e6a2e86-1dfb-49eb-bca1-21e2781a527e · inbound

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning cites this paper.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.316238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.316238Z digest=sha256:d0371576b71f1d841f2fe704d270bb9b67ebc7dcd7d0792eb5c3d8c715124775

Observation 574b4f9d-c622-4a9a-873a-34bff906c65b · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:41.022429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:41.022429Z digest=sha256:c3a6f9f9bf44373ea97415fda924ed080435ee55dea0c768d58db563589dc6db

Observation 8423a4ca-4720-4e34-9c9a-ca8ba783decd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.826882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.826882Z digest=sha256:06cc873e31c88636ae5f6ab062b5d246f2042d928a18d183cc208629811c89a6

Observation 916ad6f6-f012-41d2-8414-a9fbb17f21ee · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.069605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.069605Z digest=sha256:4a6950308d3b0dae0f3032d8e048a91c7d950f2681dfade1748521cef45313ac

Observation e7c2c932-e692-4384-ad68-e40230d919db · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.960696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.960696Z digest=sha256:3f3ff0618da8d031ddbd58392288d5d6059882f7a14184ab89ae5c9eb8de2cd9

Observation 6c208afa-7b3a-4326-84bd-828646d8ed1c · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:36.832306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:36.832306Z digest=sha256:0b3c97d4da0bd426effb7b2263df6c29eb3c8468aefdd7b460eb154f85baf614

Observation 5599eacc-8ae3-4e23-9f71-3f74990be721 · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.949136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.949136Z digest=sha256:18a40fedd92819c576f5efdaf8150fc14493aafebec354301ed1ec3676a1a993

Observation d1737b1d-a775-4801-93c3-1501881e726a · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:36:47.456783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:36:23.882344Z digest=sha256:a7f64ecf8129892603e1a86627a673c5ab26f480ccb61837a894b74cbb32401b

Observation 4b329b83-cbf4-4dc4-b7f5-9c8b8ee0aaee · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:15.882098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:15.882098Z digest=sha256:1613bb23c7f6b91c3b81caee023bf5ee0b2f466286a92d12bb1bb56c66b2e230

Observation a9274f2a-95f3-4387-aa55-e4ea777a042f · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.772074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.772074Z digest=sha256:a07e60de09e4b850a6e519429c913806d8320f2413a095a3d4beaf74c32c81f4

Observation a3597325-4084-4808-804f-767418eaaee7 · inbound

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs cites this paper.

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T06:54:11.503514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:54:11.503514Z digest=sha256:736710a0bd7e4d9631b28569962cc77bee03a3737763a3f0bf0f1d3599f832eb

Observation 30a2df69-b7c7-4aa0-bc5f-0db29504e29f · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.200106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:49d29859c5e70a355b5c4813ee8e85ab57bb8a0706ae7b0dc9420c524e808c11

Observation 539c423a-598c-47bd-a544-9f2cdb64d3e4 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.626971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:37c02f790c794a9cdb9cd99db849be36787ce05f3b2c209a06fb7d13e8ebf7cd

Observation 4a20819c-bd6d-4dac-97c7-d940b96d1f11 · inbound

Continual Safety Alignment via Gradient-Based Sample Selection cites this paper.

Continual Safety Alignment via Gradient-Based Sample Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:16:54.723710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:16:53.472918Z digest=sha256:853d3cd4c624e4e80ea0fb4f0a0f098b321bfe11ae039312e252cf50bba2ee9b

Observation 3982e68c-4587-4aa8-b7ef-d5b29f6f0bb2 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.395425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:fe2375722806630a7440d5fe58ca9ea4de50d6e95ea444e8f1362726a2b0135c

Observation ac357e15-b452-4ffa-9014-d6ea0e5def20 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.500341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:cd4e26e337f9ca6430e8d34f8cd1037d8f38d433bc0e4533807fea35a50e8325

Observation 6c007f23-3612-4364-9134-03e5d47b5c66 · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.005373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:f29e48c15ab7a9ed3b0a19b80065742493b4b0c12a908aa5fdf5ccd9a7c6d19d

Observation 026f3de0-b29c-4319-8c5b-8b46131ae44a · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:49.280147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:ad4007162d6946a68330db61dd7e4c3b66f25e0a10c30200eadd0d9979f57836

Observation 6b05f4bd-9405-42b9-bc68-ec7f5a25c467 · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:07:27.049638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:25fabd6603a1fcd765a027c28f2d29b3d91c017153714eb85c2fc60958e79c1f

Observation ca6d5dbd-88cc-414b-a563-b525b536752d · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.535416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:f74cda25ed9148751a16dfd2a472bd2195265187212e4912f44e3bac819d7369

Observation 363a98b2-587f-40db-bea0-4c105b6ccaa9 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.065497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:09718c5a076393d54e5428b9dec6e7e7ad524cf4fa36a76a5009b9a0a1395de4

Observation f2e55edf-6079-4b2a-bbfc-0111bee5acf3 · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:08:50.464701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:f80fe19e85a8c5d59830406ad0f2ffed1be8bd88682571e514170f3bb2982a7f

Observation 0b84fdc6-85e8-4316-85e1-129d7639b7ab · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:36:21.003990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:d89fd06475511d2c01347b3b834201c60890ee610e2fa36e4e9505fca5f9df70

Observation edb54687-7a19-4cea-a398-f22302efe7ad · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T16:04:52.548154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:5d29e0c2d07c11ebcec5927005755d453498d9be2eb5b09edcebf8fe355e7367

Observation 7cacd102-5732-4aed-8bc9-feaef39b7a1c · inbound

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks cites this paper.

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:53.646667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T19:23:47.574214Z digest=sha256:907b501e8eabf8de9dfb0cc42b1383b5ee2ad72574f81ddfdd26f2b6d4a6c9c4

Observation 42583a00-abd0-4214-a5eb-7b6dfc4c5c72 · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.660121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:76d220f63a021ef542865437313d0359303eb6672fadcef7a68610a596110885

Observation 86a45327-cef7-4f8f-bcae-13b1650f33df · inbound

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models cites this paper.

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.777900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:48:36.304776Z digest=sha256:c619d9a00d0ac3620c7338a21f23c8c93ad50c17b937847065f88b301a92c0fd

Observation d0ffbb66-00a2-4a1e-8735-f2177e4e3f6b · inbound

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization cites this paper.

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:13.494789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:41:03.219581Z digest=sha256:e52de64caae0925e028217aae7bd6ed8c8156303b7762d38af20ebe7304922b7

Observation bf2da3a5-6d73-4855-ad5b-43ac28e24aad · inbound

CSULoRA: Closest Safe Update Low-Rank Adaptation cites this paper.

CSULoRA: Closest Safe Update Low-Rank Adaptation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.752828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:25:01.052975Z digest=sha256:73f2fa61c8fd39415af759f4d0377d7b307533377b19bd44e1514f5fa8993a5d

Observation 916ab45f-0e62-42bf-937a-b2dac39d344a · inbound

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning cites this paper.

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:46:10.220470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:09:02.498712Z digest=sha256:6bf66f0f52fcfc0edf2198b262051286fe16df2beabc5f9396ca444d8da67cc3

Observation c1a7d728-736d-4eb8-b6dc-091baa690fe3 · inbound

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models cites this paper.

CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:18.023666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:22:01.257829Z digest=sha256:b75aef91fc2e10f5d823c2bb86585d12b5165baec02fdf38933c521738c7c2f3

Observation 0d5f89a3-ee3c-4a53-9222-020f6c0cd61a · inbound

Jailbreaking Multimodal Large Language Models using Multi-Clip Video cites this paper.

Jailbreaking Multimodal Large Language Models using Multi-Clip Video Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:36:17.122384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:16:48.957645Z digest=sha256:21ad60808f00ae47d507f230723011fc6599f25565ba384882867f165ed5ffae

Observation f7687dd7-8ec9-48f2-ad93-d68d17261cf4 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.889072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:b4854c9241e7734033799baf5a08de1e38deef40b3beceb2e9146e9b534f2e5b

Observation d5d76cf4-66d7-428d-b646-4fc822d144a0 · inbound

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts cites this paper.

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:57:38.392850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:32:18.368158Z digest=sha256:ebecfafc9ce340b021be29c2d16d83ddc9ae2ff6fe5267c61de34c1d2bf8b5f3

Observation 6b09bb8b-3d4c-4d57-b297-7c818bb3df68 · inbound

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing cites this paper.

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.364793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:02:21.293918Z digest=sha256:8a83a095fa2ecdda6b566d692c87762252fff956890dddc0f990cbda21e1b699

Observation 75df2934-0219-4bd2-aefc-b5a00ad87e95 · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.686325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:d6699b7e113454f07cd3a9615e60bd6655425def65e2223408b649e748c37e9e

Observation 0be25dfd-52ef-48ac-a9fa-a35c7f8ac810 · inbound

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection cites this paper.

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:18.893870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T20:40:15.506976Z digest=sha256:45bcc96f4b689638f8972bff9e1f0de7e8165e6af1d0c80bcb7f29f73879dba5

Observation ff6e584a-c301-4845-b460-02cd23924670 · inbound

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations cites this paper.

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.661267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T10:23:47.981377Z digest=sha256:225aebc05e9a22ef70eb0c318918d8d78485886254b27a8c66a3fd532a022849

Observation 376d516e-062b-4eae-b132-ebc3db2dad25 · inbound

Defending Against Harmful Supervision Hidden in Benign Samples cites this paper.

Defending Against Harmful Supervision Hidden in Benign Samples Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:24:45.321547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T05:29:04.802064Z digest=sha256:a95ae7ee4da41cd62c291c84a04a7546a4ad3a4767f993af014cc223bc4a60e7

Observation 45e0734a-62f9-480b-9689-95425a50aea7 · inbound

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment cites this paper.

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.623496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:20:11.322710Z digest=sha256:50711e6f8f40bf7cd259bd4b11e456f57a8526b1f33f763f05fc980bfca52c60

Observation a23f879c-557c-4abf-a2e4-8c4812be1ddb · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:48f59297facfffd09fc2619eeb667ff88cc53fee37356d9b804bf2a1809468dd

Observation 6271f3ad-1a30-4338-bb64-853b857090af · inbound

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment cites this paper.

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:10.311613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:10.311613Z digest=sha256:4f70d7653999cb0cacf56b0c8c4a8d30b6757476aa499609a8e4ad6b32c26c6c

Observation 6dc0aa6b-640a-4bd9-8fbc-554c13c04373 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:12.284016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:12.284016Z digest=sha256:709897d6b84e0482266ed2069645f2194acaf953878056fbd8fd88ae87dc4245