Pith. sign in

Paper Citation Record · LEDGER

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning

As of 23 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.00676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00676 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:34.839326Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de9f9c04-8478-4a35-b38b-bd27df643ce5 · outbound

This paper cites an unresolved cited work.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.369370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.369370Z digest=sha256:6720b4b7b51cab294da2ee014b220d99905c9b07f3854da442274f714d76e78f

Observation 427d424f-591a-4f31-9921-131da82106cf · outbound

This paper cites Zico Kolter, and Matt Fredrikson.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Zico Kolter, and Matt Fredrikson

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:36.080670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:32.462686Z digest=sha256:a75508217183da7809d26f01eec2ac30807209829c8e175f81542c638e6bbb43

Observation f90175f8-b759-464a-b357-e13422769c53 · outbound

This paper cites Dai, and Quoc V Le.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Dai, and Quoc V Le

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.576304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.576304Z digest=sha256:b3a5daea4ea7c72006db3fe74e263f12e827ec1a848e2d27bffad461a2335585

Observation 0df60811-3b14-484c-b131-d331a39ac21d · outbound

This paper cites Training language models to follow instructions with human feedback.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Training language models to follow instructions with human feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.728236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.728236Z digest=sha256:e754b63b4ba2366a7ccf03ee84125010b0aee034d032c66bcf07a560dc4b883f

Observation 4ef84feb-83b5-440b-b385-26ebd177d1e8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Direct preference optimization: Your language model is secretly a reward model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.818877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.818877Z digest=sha256:357d746eb7b8c382e5421f667c4b0efdd5250c8e450194b0f211ea82faaf8512

Observation 2b992868-0ff3-44c0-b7b2-7794e294ae1a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.872161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.872161Z digest=sha256:0ebcef692d89d2d672403087273888b0aeb4601dcdab001c0cdd588d58d64bce

Observation 5b6e2760-ce4e-47db-9af0-a46a4b894e1a · outbound

This paper cites an unresolved cited work.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:36.017467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:32.934431Z digest=sha256:1b6d88876876f455782e8d6fb9ff6c27b365a755d4cd71c496f21ed09bed5715

Observation c4aa1255-316b-49bd-9c0f-f8014a9a093d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Gemini: A Family of Highly Capable Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.985402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.985402Z digest=sha256:831fd4766404b03eee296b2b6873d1cdcb08e1b420984de5ef2df829bea563d8

Observation 8713b269-1393-43b7-a7ca-55333114f870 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Lora: Low-rank adaptation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.040776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.040776Z digest=sha256:f31c2ef63955780c59b6a88d4318327bd66f5d25f3a02b938255ae40c732855e

Observation f634a9af-af2a-4454-bc05-8e043d5ce904 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.116622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.116622Z digest=sha256:b9a371b8ae7dcd2b3511f3e79f3810dd7a29f29a19474660e1850e0bf9a47a30

Observation 960ad58d-a2de-4693-9111-0cb30b55e1e7 · outbound

This paper cites Pissa: Principal singular values and singular vectors adaptation of large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Pissa: Principal singular values and singular vectors adaptation of large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.171297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.171297Z digest=sha256:a95524dc84355b9643292cb96d026e1720cc10210a816c8711049d0806623da6

Observation 7be6ac48-e3bd-41a8-bc43-8a79b5adbce5 · outbound

This paper cites Badam: A memory efficient full parameter optimization method for large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Badam: A memory efficient full parameter optimization method for large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.978076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.268272Z digest=sha256:df0aceb2a124be2d3f3da2f296d48e85f5c3d3714defbb1caac12c28f09b4eae

Observation 7f30c869-55e7-4b8e-95a4-2b6cf3286f8c · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Visual adversarial examples jailbreak aligned large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.960806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.319627Z digest=sha256:05d98a2696113c844d13d2b1560f99fb07550552fa4bf02224182ca4450deddd

Observation 19e501af-73cf-4e17-9410-717948bfc92c · outbound

This paper cites Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.934984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.369132Z digest=sha256:ac1aa167c6bff63251e88ec93fd841c6b73b130990e9b96e4a945820517bd647

Observation 89dc442d-4a46-413a-bad7-eec925eab060 · outbound

This paper cites Jailbreaking leading safety-aligned LLMs with simple adaptive attacks.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Jailbreaking leading safety-aligned LLMs with simple adaptive attacks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.909348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.407286Z digest=sha256:694107a5ddc19c10938f2cd1b24241798b9c28c8fea752ba5169d4d343f032c1

Observation 9ff44d53-0ff6-4a37-a77c-cb02b38060c7 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.888574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.490650Z digest=sha256:1eefa1c31339c0717069dd0b94fb2e8cce8ec9892695e45e266440d002f6ce2c

Observation 070c56b1-b024-4089-b0af-2b10a0bbeb88 · outbound

This paper cites Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.866683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.549281Z digest=sha256:f1bb78646668cdc09276a4648ebb8f3ec3e70a80e0a1e824a9bccd0fc2dfa94e

Observation 15a4efd1-f46c-4d61-8bf4-36c3999ba478 · outbound

This paper cites Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.841077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.627056Z digest=sha256:d4053b58a547033854a7e1e00b25f5e18b3b13ef09fde6280df8e0fbdb2019bf

Observation 4950454c-2309-4f59-b35d-9bdb1fb441c3 · outbound

This paper cites Keeping LLMs aligned after fine-tuning: The crucial role of prompt templates.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Keeping LLMs aligned after fine-tuning: The crucial role of prompt templates

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.815897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.705062Z digest=sha256:a1dc36ebd9b3cbb812b3b33247f84953f818dbda4c00e21d62c5002b2e2298d1

Observation 05d67407-e3c4-4ece-a8dd-92fa0b9dec88 · outbound

This paper cites Safe loRA: The silver lining of reducing safety risks when finetuning large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safe loRA: The silver lining of reducing safety risks when finetuning large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.796442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.758826Z digest=sha256:6aab9226ff7fda72fd7ba07a77722b682afdadcafeeffad13dc373cfccd2e879

Observation a6f6a29c-114a-4b38-8b5c-51670c8fb2c7 · outbound

This paper cites Tamper-resistant safeguards for open-weight LLMs.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Tamper-resistant safeguards for open-weight LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.778263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.806174Z digest=sha256:d17e716cbd0ba4a1385c7ce3dfd0d5a1373274c500551511969258330f6fbd9e

Observation 66edb223-d40c-4319-8cb9-b637a74fb4f7 · outbound

This paper cites SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.759654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:33.917355Z digest=sha256:8e0a1b32440e8f80cf02bef514525efa74a22d28df3362431c019e57bc345625

Observation 532f6c18-5cea-401a-85e0-cf29774fab08 · outbound

This paper cites Safety layers in aligned large language models: The key to LLM security.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety layers in aligned large language models: The key to LLM security

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.998234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.998234Z digest=sha256:6c12629b8709b6e5ede5762cf0e30de75b3badb3dc62636aad3100bf9889e613

Observation 33abe46c-c251-4133-85d3-d9ab564bb628 · outbound

This paper cites SaloRA: Safety- alignment preserved low-rank adaptation.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SaloRA: Safety- alignment preserved low-rank adaptation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.730166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.037037Z digest=sha256:ae4a394be49f6348b9744532a486323d7e14164eb1b9e0b3928c31a5b6661f6d

Observation be3634a1-8495-4f10-b0bd-bc639ee0f10d · outbound

This paper cites Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.709450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.060960Z digest=sha256:c075267cefaa8e3b1c9ffe767c750a737e60030ee83bba6446c5a4103c005ce2

Observation 32b908e5-7d3c-4586-a750-fd4df42093f9 · outbound

This paper cites Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.692320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.090682Z digest=sha256:05065afbf8c42e7e55447388661b879e8f99b377a5c08696a149a2d2f1bf65cc

Observation dd8fa39e-d18d-4541-b2f9-f35aeb625878 · outbound

This paper cites Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.675398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.194340Z digest=sha256:8a543a11564a74281231e4cfeee3f2b1364b8ca859a6048cc7d56e7a7a561fbf

Observation 5920d55a-850f-44df-8058-206b8b396850 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety alignment should be made more than just a few tokens deep

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.257702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.257702Z digest=sha256:f75d03b88bd22958b344a7304d1e1f98d8b2b6ac3d230c8c1a18e7e6f1a28700

Observation 063617ce-60fb-4340-ac7a-71aaf5b3622b · outbound

This paper cites Defending against unforeseen failure modes with latent adversarial training, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Defending against unforeseen failure modes with latent adversarial training, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.644201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.333322Z digest=sha256:4ad40a275196498f9e6f5543606ac7e6da2f37d586c737af7f5b7149eba3ab70

Observation 479a383f-5ceb-4520-ae5f-c557d2eecf89 · outbound

This paper cites Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.627831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.410064Z digest=sha256:9c55088bba941a231dff49270356fd849b4c962ed6c707ba0c6c12353549d636

Observation 38cbbcc2-9445-4174-925c-36f6299ec164 · outbound

This paper cites Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation, 2025.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.601893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.478655Z digest=sha256:685529a4f84e039c99bfa62c7d0c4c9f2024022598af848aeca61bbbca8f2802

Observation 36b4a338-cdeb-46ba-8666-ca241a3f0cd2 · outbound

This paper cites Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.578472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.563306Z digest=sha256:137492eba17af2181eb988501cd01c68003154a36d75ced4a0cfb5a1e1f92227

Observation 0220af5e-d502-47e3-8036-525afa5ab5d0 · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.676929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.676929Z digest=sha256:a952463bfc3921ba6374aed2fdc58c4f37e5bc0d90ad88f930ff2e11b63c8051

Observation 98dead1d-25fa-4373-9a91-24834ff66963 · outbound

This paper cites an unresolved cited work.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:35.550215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.709566Z digest=sha256:815a87087ad09810bd04a41a4980fd4d29add636c3dec2aa022909cb6b057a92

Observation fc923ced-3d93-49ee-b673-1ffdbd439a66 · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifications.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Assessing the brittleness of safety alignment via pruning and low-rank modifications

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.524485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.714063Z digest=sha256:3721e2d58b5c682fc90b8dd8ab81de724817b6abebae7cb64b3fd7d6291312f3

Observation 7f151a94-4360-40f2-b2d0-2976dc036c4b · outbound

This paper cites Toward secure tuning: Mitigating security risks from instruction fine-tuning, 2025.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Toward secure tuning: Mitigating security risks from instruction fine-tuning, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.505453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.718187Z digest=sha256:7d298690e0d896226581a67cac661300fa0c1b4defac04a15c2d76410ed327bd

Observation 7f17451b-dd86-4a84-830c-b9a2252ea844 · outbound

This paper cites Locking down the finetuned llms safety, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Locking down the finetuned llms safety, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.486382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.721988Z digest=sha256:a0c320178c2c127c6f8545c2675e503e72d21321879db3cf7f538a6a8a45299e

Observation c6d5ad28-eef6-4875-94b5-caa73e79424c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.726375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.726375Z digest=sha256:8a0c40563353cfc0b5346fc0574ff810fe7a4cc2ae43bea2a721569ff31be3e9

Observation 4e2e8df3-365b-4eaa-b24b-c22f448e7c63 · outbound

This paper cites Measuring massive multitask language understanding, 2021.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Measuring massive multitask language understanding, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.731783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.731783Z digest=sha256:89423ce87d1f7a0864e5bacf9df8327f5d27c70b2e421dc9b54bd9d430da92af

Observation 4edab1b0-6c34-4396-9b13-872213a50348 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Xing, Hao Zhang, Joseph E

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.738615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.738615Z digest=sha256:42ebfeaf112a004731e9993393583c16f42dcea770ee8136fe5a67da3cc4ea1a

Observation 1bbfd0b5-32a2-4dff-95fd-f293d9622854 · outbound

This paper cites Certified defenses for data poisoning attacks.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Certified defenses for data poisoning attacks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.428950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.743466Z digest=sha256:a2f2608407075a9117f2dc891b221acb6d787a9328459e6d907928cc2c4d2cce

Observation 61c38025-b471-486b-b429-ff71b0240bfa · outbound

This paper cites Data poisoning attacks against federated learning systems.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Data poisoning attacks against federated learning systems

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.394671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.748082Z digest=sha256:2fb99363bdf73e7a799edb46b28cf894c66b3317f54578874ed483a81cbd2fc2

Observation c26f605a-8aa8-4e12-ad5c-382f47996d28 · outbound

This paper cites Immunization against harmful fine-tuning attacks.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Immunization against harmful fine-tuning attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.757469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.757469Z digest=sha256:f7803815a6f1c8591a45948079e64e44a514eab65b6f752e09edbb7a83d3e881

Observation 77a852dd-4a36-4ce8-a082-12588c67ed17 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.761861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.761861Z digest=sha256:89648fbda32b167a860b0ca1cd023c0e30e1bf3ca67a111af54a7bde6779201b

Observation c7ec8314-dc2d-4267-8e4f-5a0ad634537f · outbound

This paper cites The case for an opinionated, theory-oriented real-time operating system.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning The case for an opinionated, theory-oriented real-time operating system

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.372479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.766914Z digest=sha256:9ed1e80211df07ea0e5d48c0e19ec035c1b684c29157fe768644ee8275d4e27f

Observation 0a34b9a9-a42a-4873-81c9-41aee11e8d93 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.771964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.771964Z digest=sha256:8caf60f1ba95758ea5053aa1438a01e693fc80d220876c32e53065e8c0a2ef6b

Observation c324371c-2ac5-4ddb-b0ec-8e653cc155b7 · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Recursive deep models for semantic compositionality over a sentiment treebank

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.776353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.776353Z digest=sha256:119d328348b2ce641f43ae5f465c407790f80b55615a0995ff5a8b86a70bc0bf

Observation 9ca2bf94-8217-4af6-8aa8-00a1a4362ebb · outbound

This paper cites Character-level convolutional networks for text classification.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Character-level convolutional networks for text classification

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.277370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.780136Z digest=sha256:290f3a67d3665fe78dd57105743cccf7c567c966da06a3cd2b2724381437d257

Observation d4746240-e2a8-4e3f-947b-886a93bb3f2f · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.784487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.784487Z digest=sha256:0002574e6af096aa43fa6302a928fd5e5265c42ac41ea5aa97427b7a481ab820

Observation 5f838a6b-4c60-4c9b-90b8-d95bb9dde953 · outbound

This paper cites Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.789925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.789925Z digest=sha256:17eca0b2f55db2624ab6b0d6a4d0d8a0da660f9ae2c96d89cc337d3d85b5daf4

Observation 4890fa07-bb75-4ae3-8670-a364756f90c1 · outbound

This paper cites Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.795389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.795389Z digest=sha256:8c3a0d92c574fe13a5ebaabac603c32f037d4d8e0b5ca40c47016ef7b7c1d481

Observation b2893360-00ca-464f-b0f3-8d2aafcc35fe · outbound

This paper cites Alpaca: A strong, replicable instruction- following model.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Alpaca: A strong, replicable instruction- following model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.219475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.800016Z digest=sha256:cf925ec80485baaf755d39e27e2dff5566e479d2416afda0fe919b82f8009ce6

Observation 463654fc-4125-455b-a554-7367748f69e6 · outbound

This paper cites Introducing the world’s first truly open instruction-tuned llm.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Introducing the world’s first truly open instruction-tuned llm

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.201334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.804830Z digest=sha256:6ade1b71ccbcdf8367f5eae6261d98ac3a2843fa9fbfe7d67b36757619b2f4fe

Observation 0e5921d9-feef-43d2-bd67-43c267879912 · outbound

This paper cites q_proj",.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning q_proj",

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.174645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.809340Z digest=sha256:ae5ee2cf4d818e9adf4f1a22eaa170fbebd963878b234b8393112e4f4e4bfe19

Observation 498903de-128a-47f0-8494-6b248ff36c3f · outbound

This paper cites This is the most preferred response when confronted with policy-violating instructions.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning This is the most preferred response when confronted with policy-violating instructions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.815792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.815792Z digest=sha256:da4c936dc060fae10e27d349d13a6d8f46efe2da399505f90020cc755c2fcef2

Observation feb12fda-29ea-44dc-a028-3c21b644ba91 · outbound

This paper cites It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.821541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.821541Z digest=sha256:0f37b9c86af6b260c5ffabd9fece0d63fbe4ed18493fd3418b683bee49c569aa

Observation e23d4b7a-4399-4401-96c0-e9c832dcc3c1 · outbound

This paper cites It offers content related to the request but without embedding necessary precautions or disclaimers.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning It offers content related to the request but without embedding necessary precautions or disclaimers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.827143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.827143Z digest=sha256:b47b779932cfc9498400ab6f70fe35c62020d93b0548505bb9cc81a88ea2e63e

Observation 74b6152a-c0c0-4c9c-b8c3-eff51f401bda · outbound

This paper cites However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.832822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.832822Z digest=sha256:b44cf4aef6dcab0f21db6e2f5a19cd9686dac50a05de1ff064f9633478548851

Observation 0f610a2d-f3a6-44ba-99ce-8512d95179bc · outbound

This paper cites Candidate X Under Investigation for Embezzlement.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Candidate X Under Investigation for Embezzlement

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.072747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:02:34.839326Z digest=sha256:a225b5ad65da41efc4c83d23ef388afa146fe841ae198aea7e1aedaf9b6c08e3

Pith citing papers

No inbound Pith citation observations are available.