Pith. sign in

Paper Citation Record · LEDGER

LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2310.20624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20624 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.171392Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.050459Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9316efe9-cb13-4923-b761-37516f01734b · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.010904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:50643eb1dfb17b0d2951fdea5885b9c39e15fab4fea13206bdf2f4b9f3a8435f

Observation 6e7dbe39-3217-4a9a-8f83-2d7baaabdc48 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.710807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:e8539d9047b58e530fd42153a100df67f9f53b09120e3ce29535d3d84f573d28

Observation 512ac141-c77c-4a46-bde9-1738150efb2a · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.335120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:b471c500c4c2bf7444d13966e57a2af73e7a6d2a45ee73f38822942c73027e60

Observation 43acf8e1-359d-47e2-b21a-ca64edf4fd8d · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.171392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.171392Z digest=sha256:7a80ad9e66e8478081f6e13c1d14cbc3a252c68cd9f20e1858a9af0fcb50990c

Observation c79286d2-e8bf-4dca-bc5e-22518a51632a · inbound

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models cites this paper.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.390992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.390992Z digest=sha256:9e6265d1d83bfe7fc3289f9cbe79407430bd390ea9718a89f7a475d16983d7b7

Observation 001670e7-7c84-4315-bf67-bd245bd8a832 · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.267212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.267212Z digest=sha256:622f1bc856e040590ed83a0ec65c409e07a6b7a2b6240a0a285e7274026b0724

Observation ad933c1b-7a01-4ba1-8ed3-cb314ddb6a47 · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:38.516672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:38.516672Z digest=sha256:1bbb57e4997371bf546eb0f6616db2d82bf886f201b39cb3c19237b91379c364

Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · inbound

Linearly Decoding Refused Knowledge in Aligned Language Models cites this paper.

Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.939664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.939664Z digest=sha256:49060ad40f8cedf9a659a3f3d9f59374ed2ee52afe9c367eb77a5540ecf77b97

Observation 111021dd-05f6-465a-b48e-2376af7f3472 · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.851271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.851271Z digest=sha256:bf5e7978313a4e9e47a9eeb4e554cadd3d2d53ddc868562ab33610d458c8a026

Observation 30f57e83-21f8-4faa-998b-68177e063140 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.084303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.084303Z digest=sha256:4a3daf59555088ba163e2507cfd11ff73ca21fd32f689a3d858f4ca90cb34b89

Observation 42d9129f-7bb4-4215-ae7b-53ab8a841f8e · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.299965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.299965Z digest=sha256:8a4267dd7c4164366b0f55c42f613c4541d4eae34dff68667663101e6be4c801

Observation b9f2add8-82a1-40b4-ae54-e40684a089a7 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.430368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:74837c0fd40070d2d7725ffd984a9bdc6b71c461d5c0452498cd8ee2a5c6845b

Observation 755422c7-39cb-4e43-a9dd-d8e7670c5380 · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.850194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.850194Z digest=sha256:64ca351d481b4459837cb4d05be9ab46c7fbfec90dc35177b24faba1699c12c4

Observation f0230405-d481-4232-a71e-a6bd91d0f21a · inbound

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models cites this paper.

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:27.753084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:27.753084Z digest=sha256:bfd394862d2698f585f31fb9ed3cec187c7f3dd529072d913ab11d83ce051d31

Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.450991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.450991Z digest=sha256:e410f80fc0f5f874de477f5141880b324e6dac139409c8f7d9a02d3ee71f5570

Observation 3407591e-bc9f-4f12-bf67-eb608e0b06f3 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.907738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.907738Z digest=sha256:b7ad3afc97d4d1ff1245b56eb7236267286153261d9d74fbdc9a74f2c1994d45

Observation f39bbd99-254f-4599-8cfa-917f2b9d2688 · inbound

Understanding the Effects of Safety Unalignment on Large Language Models cites this paper.

Understanding the Effects of Safety Unalignment on Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:38:14.463635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T20:37:45.061026Z digest=sha256:b04e0c3597cea0383397eb61ca76c53f1688b9707f512918e8e8a34cb3067944

Observation 51b9dee9-9552-4825-aa11-e8f2e9bbb3a6 · inbound

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs cites this paper.

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.899967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:01:25.938248Z digest=sha256:dff8e5e8627e78c289fccde72f8e5d137109a50654e884e0f3255c86f0b6a6ba

Observation 422f1b45-3cb3-4390-9e7c-d6df5bc3ef48 · inbound

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks cites this paper.

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.698001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T04:24:25.964495Z digest=sha256:ca57dfe19f5eb85b129dd8ee8d1ef48ab9130671aceb598bee2fa3e1c77f29ac

Observation da4a9885-859d-4dbf-8363-6f186afdc9e5 · inbound

Low-Rank Adaptation Redux for Large Models cites this paper.

Low-Rank Adaptation Redux for Large Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:26:03.822655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T21:48:48.992712Z digest=sha256:f5e591ddfcbec384e97a0c73a28a7a397c09dfa200e1f6bb2a9bdf3bc5a0cf96

Observation 8a03f853-c1c6-4675-89be-89d7d32986a4 · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:19.248358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:344bad162e29405ec08b18fcbc050f780be308ca16a4aee30d6bd6869e9642cf

Observation 0fc8591f-5942-4348-ae96-eb953144eb4b · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:08.674491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:22c18e19a74a58cdfa2ff3ce46a0377d5b2a4021fddf6ec2b9266ba404681378

Observation 2d233b8b-046f-482e-86fd-c6c11a5a36ed · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:07:27.074141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:9d5f34b3ba035f7b51ef58b75b95a674e4a988ac75d7d61b869144dc25cc9d18

Observation d80882a7-057f-4b7e-8b5e-9fb5cb69ae57 · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.752995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:44:09.214779Z digest=sha256:f85773e18ace3520b5a0461ed761a87eb63a7d42affd09aa2cbb39dea5bb9d5a

Observation 006257a1-00c8-49c0-bb49-ba8b76145c5f · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.617390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:05:29.444682Z digest=sha256:804da21ade916589128bcdd01932604457386c095f7b4437602f0c088da41b99

Observation f51e1717-a090-434d-a1fa-d80497739ad6 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.074611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:450acd5129ca77a59c96ff71da57ff012452b2e13aa567fbc8e06f80dbe578bc

Observation 7b52221c-a0a7-4743-9217-5e493c2088d3 · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:36:21.008543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:bd1ea133c935239b9f1b5e28053b79714de597415911f3bf39d7d39c13b5f867

Observation a77825fa-713f-43ea-b688-d21710ce826e · inbound

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models cites this paper.

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.288540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T12:58:50.713928Z digest=sha256:c1f41a2dbbb6768c21d6cb753675ff97b936e61084ff611a5bb7b570efced422

Observation fb7efac4-ad6d-4700-806d-b596f98a71b9 · inbound

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol cites this paper.

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:14:45.888901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T14:05:18.213065Z digest=sha256:79b6661dbe86b8cb670627bf454768d14b5a1997d7b634a054e0914142b49e24

Observation 76c0ffc5-9fb8-420a-bd4e-450a48bf5004 · inbound

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions cites this paper.

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:54:03.326066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T23:51:05.413122Z digest=sha256:110dc8a846f8e2a539768a7338cd73198d45fd6d523190e489bd9429f8947c6f

Observation e8bf7610-d2e3-4ecc-a0ea-df10c8c82bca · inbound

Curriculum Learning for Safety Alignment cites this paper.

Curriculum Learning for Safety Alignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.186396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:25:31.739331Z digest=sha256:9a92150bb96b4a5de042dc671bae74c2cf77da748bed6a88f01c1024ca47fda1

Observation 671d674a-6cd5-42cc-b508-bc3daa08761e · inbound

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks cites this paper.

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:53.626743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T19:23:47.574214Z digest=sha256:405bf13272ab19cb68c64d4c52245e427d2139ae8431d492e0bc87160d801f58

Observation 4fe3cc67-1891-4f66-ace7-8f8b60f7b0ee · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.681269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:a7d0f26745760e7a80b193bf73e7f8bb45e54b3202a7577e4d13d74be70dc1f0

Observation 29586d0c-7565-48ab-9965-d676f03aa345 · inbound

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning cites this paper.

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.225624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:09:02.498712Z digest=sha256:265294f9e00933b9a2ba01468b49411b564a60a1940f419e58f9f32bf4193e5d

Observation 93a5ac7b-9a1f-40a6-af55-c8c43f8b5b4a · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.872764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:6a1d52fb71cf6e5a220ec9089a6a5c87f0b9b17777546124d634a8ce05df8188

Observation d1f12f49-34ff-415b-8778-47ce65e5716a · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.234243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:504da2f525d74665c883f3f2ad117a25a24690ec832927073df2309f2e982921

Observation 0d16d39a-46d0-490c-bbfe-bfd4c1cd6bda · inbound

Discovering Millions of Interpretable Features with Sparse Autoencoders cites this paper.

Discovering Millions of Interpretable Features with Sparse Autoencoders LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.052072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T05:34:00.754172Z digest=sha256:a1729c7c4bfcd1f0d62437e687a76fc73545d9957eb76b58834a8cb7b0a497b9

Observation 7763be8f-401c-46ff-834f-475b65181e5f · inbound

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning cites this paper.

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:34.733665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T09:52:42.862423Z digest=sha256:27e34ba7f540073993925b3aa9b1ad54e14fceefaffe9b54b8dc0a9eef951960

Observation 150f3f6e-8b64-4988-9ed7-8e975ea53f56 · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:5d2ece46998363ed251bfc71ff171f2182cd22e6e12be8c77099240814af398b

Observation a946a932-8846-47e0-af95-2f39d7873f13 · inbound

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment cites this paper.

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:11.935152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:11.935152Z digest=sha256:0deec4da3abbafed7ea9a07c1f7a34fb505236ee925863cbb99a7864e350f329

Observation ed4e70e1-67fb-424a-937f-79b556eadb57 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:09.563432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:09.563432Z digest=sha256:1eae778be49bdb9ed40d24700ab7f0ccb8efb3fd3f80f34a3fd7bed82bd301f5