Pith. sign in

Paper Citation Record · LEDGER

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning

As of 13 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2412.12497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12497 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:06:48.229957Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40854e01-e74b-4965-b0d9-cbe40d469c66 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.058083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.058083Z digest=sha256:a82076229f35ebe7b4ec7fcbb35b979051db1845356d32c292c135bfd4640a49

Observation f62d56ff-3c3c-462a-a48d-5c6e327e6c3f · outbound

This paper cites write newline.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.063975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.063975Z digest=sha256:62761a37b693beb3b7971d36bab10e113f1ae51e85dd8b3171ea1f156f61e99a

Observation 97e7f11b-9295-469f-b56e-e7c105f557f3 · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.069054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.069054Z digest=sha256:2b91c60b20892df26b7c237c7106c1a0e8c2e134cacdfee5c1921e1edef147dc

Observation 8fe5e615-492f-4fec-9fef-51ca9752eaf5 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.075163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.075163Z digest=sha256:ab3aa02e45e8206a9fd9f72dea2347367f382da737bb0b393707b373324e559a

Observation b08d789a-be3f-417c-8a36-623d179d9365 · outbound

This paper cites F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:48.796303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.079786Z digest=sha256:c3dc5f8bb31397967a3e6cc9d0993f099d587df9fbab470e09d69ebaaadd70e5

Observation d9decf9f-77b8-401e-97d0-1fa8f5e22ec8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.084434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.084434Z digest=sha256:dda7120cbe215d5037125304414e60e62b02789f2673eae1146d6fd0150982fd

Observation 6bfafd4c-6f98-4d53-b514-05e918f83dad · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.089407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.089407Z digest=sha256:71f05e411897fe9296d97906d92a7184caa72c712f409e2907774e9d810ac9df

Observation a64e96ea-2383-4421-981b-1f5701fa8642 · outbound

This paper cites DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.093168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.093168Z digest=sha256:c5f5a5edd5d73dfafe986cf4da74b672e5f1f6a30295fd8e33ef90cda4f25442

Observation dc1f3363-4a3d-41a8-bc03-4280d3fc4d82 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning KTO: Model Alignment as Prospect Theoretic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.097418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.097418Z digest=sha256:e69c86bcab76b66255e5daefa0e0ac3442f95f1920048778363e05a04fb8ae26

Observation 0239bebb-fd78-40f0-bc2f-39878c464e92 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.776979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.101760Z digest=sha256:4cd007c9d6b1417a5068b6710bfaa733241ad3c5e7fe73fa6f6d3587c85ea851

Observation 03b1b44d-4afc-4701-97e0-e3ea5d04b89e · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning ORPO: Monolithic Preference Optimization without Reference Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.105864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.105864Z digest=sha256:c5cb2dd9365c20b319a3a4dd1e33d2c3c492711aa4fc120382dae544205a808f

Observation 60247d1a-0e15-4410-b08e-c23a53876c12 · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.111897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.111897Z digest=sha256:cf543f88d34d91dbdfba64aa8fe77c9b665372905b298706595ad1d0f3716bce

Observation 3be4e412-dd8a-4cac-a863-aa41b3c84101 · outbound

This paper cites J.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning J.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:48.765676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.116771Z digest=sha256:30aedb351a946ddc02abd0f0f8e104cf9a40de7c5ad188563fca85b2ecc9f4df

Observation 1cf6fbb8-cf1c-47f2-a878-01a11a2c7bb9 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.121324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.121324Z digest=sha256:d794b8ee1772828b3647d95e282b2b962c44e178f6e5e493f10177eaa5fd8b5f

Observation 690d51f3-c27d-461a-8c21-3a61329e228b · outbound

This paper cites F.; and Liu, L.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning F.; and Liu, L

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:48.752758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.125788Z digest=sha256:f36eaab1cc23d9dc6575956496d8f3e675358a359f1dda74713e062a00c0eb81

Observation dbf4849b-520d-4b25-9590-fac183c3c6eb · outbound

This paper cites Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.130126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.130126Z digest=sha256:7ff93f49f47225d7c2c28e19dfa5ef74cd9adf5e5b6e8748f48b1cd4a2eff2ee

Observation 6365235b-d91b-493e-a612-5f4efede6104 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.134674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.134674Z digest=sha256:e97c7dbabd4d1766be643f6ff8922cc35a79bce990ddc6600d511fa359af1d8d

Observation 6bba3c2a-6f66-454b-a60c-6cccafaaa134 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.140584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.140584Z digest=sha256:9f6b2019c6fd6173f2419402ddcce76cec96a10c6d0ae0f78419986ae88ab291

Observation 34ddb55e-4d52-4603-a3ab-5c1c2d3b468e · outbound

This paper cites Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.144977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.144977Z digest=sha256:d95fce6f5ef7a07fe2e39e602a3c362df1a1c6ea403c8c6eae8a18de0a235f73

Observation 37dba6a7-9873-44ef-8f81-ea53b7fdb5a3 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.733378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.150529Z digest=sha256:9a5e2ff9fce9bb48d6614ceb9d8a86fc66f293eddb6ec2cd24ea246e8447ca58

Observation 3ada8865-7075-4e2d-ac37-d98dea9d8925 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.154835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.154835Z digest=sha256:1246d026fb2fedd240067f31f404b18632660adde03be2255d5dd684cd2d0d6a

Observation 11ee58a7-7eda-4938-b8ba-fe15bee946e9 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.720242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.159390Z digest=sha256:13382f9e714ffd1d867c1659e3c02d2239e6597fe84d15bac70baf809dfc07d2

Observation 07953f59-80f5-44a2-ae5d-c74f2578d8fc · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.163494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.163494Z digest=sha256:325e78aa3add2137b45ce5cb2f662f855c40fe8a6fcb56eab6769fad72f278aa

Observation d8a70c95-2273-4ece-b0bd-5ff11ef1c5f0 · outbound

This paper cites M.; Weber, L.; Choshen, L.; Sun, Y.; Xu, G.; and Yurochkin, M.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning M.; Weber, L.; Choshen, L.; Sun, Y.; Xu, G.; and Yurochkin, M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:48.699002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.168044Z digest=sha256:eb6dd9ec463d2091a5d15b8a6ea444e4c57485a85f2c20ea8fa68095aa5f0889

Observation 23553a85-712b-41ce-81f0-b4d386c4eac7 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.172421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.172421Z digest=sha256:2b71cdfbba49cb21c88641e01d01aae4ab3f6f398528a749d5d4c7940be197ee

Observation bebc58f5-0fab-40a1-a6c5-74861db3b877 · outbound

This paper cites Learning to Poison Large Language Models for Downstream Manipulation.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Learning to Poison Large Language Models for Downstream Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.176729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.176729Z digest=sha256:e3c501c4a61e859394893323a91001f53bdeb15f8fa55858494d7e1d9d91a41a

Observation e48056dc-b2cb-4cd1-aff1-81eae5d6b4b0 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning D.; Ermon, S.; and Finn, C

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.180673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.180673Z digest=sha256:b9d4a3fa803e847ffe29e260611dc59dea4a58da4b581e464654e0763395d846

Observation 91be9f52-0582-4a44-8897-72b6035d46a0 · outbound

This paper cites Open Problems in Technical AI Governance.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Open Problems in Technical AI Governance

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.185516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.185516Z digest=sha256:c4d930a5ac7bafb697c8c28811d7437adccbd4f25ae71547050f9cd7c6d0e548

Observation 5884aaed-a0e4-4cae-a4f4-211157209f5e · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.665244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.190088Z digest=sha256:207f5946fd1c571bac0e8c3442bce01eab8830fdf18081e5258dacb5bbd476b3

Observation 155607c6-7134-4392-92ef-3e3f3826709b · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.648895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.193565Z digest=sha256:938f2a6d125e9447b180fd435222c8c2d97762fea351bbaaa2e4dd89b36d9092

Observation 73d38436-1141-4c7c-9139-04a1373b2d4e · outbound

This paper cites D.; Ng, A.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning D.; Ng, A

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.197255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.197255Z digest=sha256:fad76ac1636f590fee07e9b397060fd9e9fe1f40effcbff41c6c4f6639d937ef

Observation 1ae85030-7255-4309-8c58-a4856bb3fc38 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.627338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.201967Z digest=sha256:8d1a7616c60a3449f8231684148487390c9f4fa331efe81edaf189e08a3e45a8

Observation b0e13af2-798e-439e-bfad-f2e1609a6e33 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.614745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.205663Z digest=sha256:222954dc42c0e3429c1ab3639415893f6e5289009705598777a2301b48f9b7a0

Observation cdbd60c0-94a6-4a42-a7d2-75353dabf277 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.601700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.209198Z digest=sha256:5ed382b0dbeb747a49abc5717d730247267d782695a99cb7efc4ef9e9975d0bb

Observation 5c404844-43ab-4d8a-8466-32a1ffca7b90 · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.213034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.213034Z digest=sha256:32a395eb81858e4d924bcd60db1bca3b917318c0efc89af7fd1530ba944f3de9

Observation 85bc9187-ca80-4bed-aeaa-c403d0630379 · outbound

This paper cites BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.217261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.217261Z digest=sha256:d7bb99b4f05ec92dde68ec0f600cad3e03bd82977d59e8d1a004761e87cd4e15

Observation 8cc23ba2-e465-4ddf-a77b-9681f0fa6f87 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.221300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.221300Z digest=sha256:6622ff8219d6512f7fa593bd9e52810be4e3856b2d5d8dd54b5cc4e2e7b0bdfb

Observation cbd7fd4a-5d4f-467b-9214-c6580f013d38 · outbound

This paper cites Model Extrapolation Expedites Alignment.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Model Extrapolation Expedites Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:48.225243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:06:48.225243Z digest=sha256:9b13ff9de9fca90f7ba865c302eb0faa40da4d80369b8fc1f319f125144f6136

Observation ddb25a55-3154-46aa-a0c3-959550278c99 · outbound

This paper cites an unresolved cited work.

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:06:48.577859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T14:06:48.229957Z digest=sha256:6afe1bcd7ab14392eadf29886ca8296626cdae4488048b3343e524c5946627e0

Pith citing papers

No inbound Pith citation observations are available.