Pith. sign in

Paper Citation Record · LEDGER

LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2310.20624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20624 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:23:02.390992Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.050459Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9316efe9-cb13-4923-b761-37516f01734b · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.010904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:e3fe3068e4cba878058dd1477dc00aeea988d4f231955bbbd7359dd5a935d53b

Observation 6e7dbe39-3217-4a9a-8f83-2d7baaabdc48 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.710807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:405024aea3c0700c6199639dfe67e92c1df2288c6806eb932cf7314eaa48609a

Observation 512ac141-c77c-4a46-bde9-1738150efb2a · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.335120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:ea1596c2c95be9f3dac7418a81f42034b79661665906768eef7b9a61c9532a89

Observation c79286d2-e8bf-4dca-bc5e-22518a51632a · inbound

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models cites this paper.

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:23:02.390992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:23:02.390992Z digest=sha256:bb3ed99fd97e0b6aa5e699a0b2ab70d6b3cb0f06f504dc02ee28f3b246945013

Observation 001670e7-7c84-4315-bf67-bd245bd8a832 · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.267212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.267212Z digest=sha256:622f1bc856e040590ed83a0ec65c409e07a6b7a2b6240a0a285e7274026b0724

Observation ad933c1b-7a01-4ba1-8ed3-cb314ddb6a47 · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:38.516672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:38.516672Z digest=sha256:1bbb57e4997371bf546eb0f6616db2d82bf886f201b39cb3c19237b91379c364

Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · inbound

Linearly Decoding Refused Knowledge in Aligned Language Models cites this paper.

Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.939664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.939664Z digest=sha256:49060ad40f8cedf9a659a3f3d9f59374ed2ee52afe9c367eb77a5540ecf77b97

Observation 111021dd-05f6-465a-b48e-2376af7f3472 · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.851271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.851271Z digest=sha256:bf5e7978313a4e9e47a9eeb4e554cadd3d2d53ddc868562ab33610d458c8a026

Observation 30f57e83-21f8-4faa-998b-68177e063140 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.084303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.084303Z digest=sha256:f417c03c6a458494b645185c7827c5da37fea825768dcc53379b050be7ea14fc

Observation 42d9129f-7bb4-4215-ae7b-53ab8a841f8e · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.299965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.299965Z digest=sha256:8a4267dd7c4164366b0f55c42f613c4541d4eae34dff68667663101e6be4c801

Observation b9f2add8-82a1-40b4-ae54-e40684a089a7 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.430368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:7b203ad34caf3be154c4ab4c0b9fb41762e0c2f4ba4a879a578ab8d4ea2d6e2a

Observation 755422c7-39cb-4e43-a9dd-d8e7670c5380 · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.850194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.850194Z digest=sha256:64ca351d481b4459837cb4d05be9ab46c7fbfec90dc35177b24faba1699c12c4

Observation f0230405-d481-4232-a71e-a6bd91d0f21a · inbound

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models cites this paper.

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:27.753084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:27.753084Z digest=sha256:b9a877c0b007647f2c27ee543765d8ed8c1284711c72ac330fe7257a28e6cd55

Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.450991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.450991Z digest=sha256:e2ddce7b68e5f48f4af44a40b4955541dafe3ba3b3c2bc1be19ddef7b1efb581

Observation 3407591e-bc9f-4f12-bf67-eb608e0b06f3 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.907738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.907738Z digest=sha256:2edc5004a0720cb98fc1a2630424ba8002828f769744daa70516a2dedc78fbcc

Observation f39bbd99-254f-4599-8cfa-917f2b9d2688 · inbound

Understanding the Effects of Safety Unalignment on Large Language Models cites this paper.

Understanding the Effects of Safety Unalignment on Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:38:14.463635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T20:37:45.061026Z digest=sha256:61e09132068ab2c264a67f427df90348e0fd8a23de55107e7215b62074a6e959

Observation 51b9dee9-9552-4825-aa11-e8f2e9bbb3a6 · inbound

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs cites this paper.

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.899967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:01:25.938248Z digest=sha256:c3bded1c8c0b7fafdfc3fe9465fd58ed24a3cab2973be5dc43c7b942c6c5d65e

Observation 422f1b45-3cb3-4390-9e7c-d6df5bc3ef48 · inbound

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks cites this paper.

Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.698001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:24:25.964495Z digest=sha256:f18d0d5b1343b42e62a5851912caf959383fbdb1415ddd204bfb89e76d2afff0

Observation da4a9885-859d-4dbf-8363-6f186afdc9e5 · inbound

Low-Rank Adaptation Redux for Large Models cites this paper.

Low-Rank Adaptation Redux for Large Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:26:03.822655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:48:48.992712Z digest=sha256:42f6017dc04b53d2fd15398cbb98135c81aae7a9dfb1e9897634bf6d0e83e3b7

Observation 8a03f853-c1c6-4675-89be-89d7d32986a4 · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:19.248358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:98d45306c7012b6571adb8eef78882a2e7b8a1274e4f5a780561d061c2f0d17e

Observation 0fc8591f-5942-4348-ae96-eb953144eb4b · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:08.674491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:c352c570256465f87b2d40329482e58d094dfcc80331276ff536cd5e66e1ad03

Observation 2d233b8b-046f-482e-86fd-c6c11a5a36ed · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:07:27.074141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:7709da6797d3d7cd9b156eec91bec376c79599f89611cd4fde20e05938006398

Observation d80882a7-057f-4b7e-8b5e-9fb5cb69ae57 · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.752995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:44:09.214779Z digest=sha256:9d997465d54618b5d0d1dac6028e061637f6d9908c35fb19f491a591eb7e771d

Observation 006257a1-00c8-49c0-bb49-ba8b76145c5f · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.617390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:05:29.444682Z digest=sha256:a307949ccb1d28b01cea4859c1790092b752281895583344721589ef30c369ec

Observation f51e1717-a090-434d-a1fa-d80497739ad6 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.074611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:dd421845e3bf960ff449a26e76f7cbdf10db9027953e359011720464870a98b5

Observation 7b52221c-a0a7-4743-9217-5e493c2088d3 · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:36:21.008543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:c1b303243915ff7d3b8bbf1bf67f96bcd287eef9f63b75b136a9ea4334d30836

Observation a77825fa-713f-43ea-b688-d21710ce826e · inbound

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models cites this paper.

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.288540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T12:58:50.713928Z digest=sha256:453caae2191be11f2e9397368f61d95369b45825423c818e4450d4cc89269e13

Observation fb7efac4-ad6d-4700-806d-b596f98a71b9 · inbound

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol cites this paper.

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:14:45.888901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T14:05:18.213065Z digest=sha256:7e8c04a6ae7fb15acdb96e15cca40693e1f3bc0759e262b26bc0342c0b6b08ef

Observation 76c0ffc5-9fb8-420a-bd4e-450a48bf5004 · inbound

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions cites this paper.

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:54:03.326066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:51:05.413122Z digest=sha256:f087d09cc8c3ec7d0c558a29d98f9c119486bbf94f5402bb7c5c60ca65e83ce6

Observation e8bf7610-d2e3-4ecc-a0ea-df10c8c82bca · inbound

Curriculum Learning for Safety Alignment cites this paper.

Curriculum Learning for Safety Alignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.186396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:25:31.739331Z digest=sha256:1579df7ff724c54883203c8dff30778c798027e35a1f6d3c0a5470119ad5011d

Observation 671d674a-6cd5-42cc-b508-bc3daa08761e · inbound

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks cites this paper.

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:53.626743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T19:23:47.574214Z digest=sha256:f0694cbe67eeebadd506c75a160fdcf5e07d523dd70cd6642c227a391caeaf86

Observation 4fe3cc67-1891-4f66-ace7-8f8b60f7b0ee · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.681269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:0ac2033206c765d16b2f12b143632d38cb7f259e48c7cec3c3548f8ed857af74

Observation 29586d0c-7565-48ab-9965-d676f03aa345 · inbound

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning cites this paper.

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.225624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:09:02.498712Z digest=sha256:11c52a9e127c0cec87d6f6325c8ed1782d900946285f92449732e26527f876e6

Observation 93a5ac7b-9a1f-40a6-af55-c8c43f8b5b4a · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.872764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:d626b96417a3ba3ab6ddc3b4999e8089f32e443fa4d5d15bdc67c59dac2b2b3c

Observation d1f12f49-34ff-415b-8778-47ce65e5716a · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.234243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:c884dcec6c08f1efc579aa0cab93f7d6795c9105f537caff45c2cdc5f86aaaf1

Observation 0d16d39a-46d0-490c-bbfe-bfd4c1cd6bda · inbound

Discovering Millions of Interpretable Features with Sparse Autoencoders cites this paper.

Discovering Millions of Interpretable Features with Sparse Autoencoders LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.052072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:34:00.754172Z digest=sha256:6e61a4fa990668924f5f7e3da45d5244140c63b24e261f84ec77eff14e3ff22b

Observation 7763be8f-401c-46ff-834f-475b65181e5f · inbound

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning cites this paper.

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:34.733665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T09:52:42.862423Z digest=sha256:f71914e3d2bc18843a6274e8c90bcab38785558af8c7ba2b5e38b7decfe836a4

Observation 150f3f6e-8b64-4988-9ed7-8e975ea53f56 · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:5d2ece46998363ed251bfc71ff171f2182cd22e6e12be8c77099240814af398b

Observation a946a932-8846-47e0-af95-2f39d7873f13 · inbound

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment cites this paper.

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:11.935152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:11.935152Z digest=sha256:0deec4da3abbafed7ea9a07c1f7a34fb505236ee925863cbb99a7864e350f329

Observation ed4e70e1-67fb-424a-937f-79b556eadb57 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:09.563432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:09.563432Z digest=sha256:c80d19254419893fd58c216c24012b718c025d4537d2364ec10bf09e57eabb03