Pith. sign in

Paper Citation Record · LEDGER

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.00676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00676 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:34.839326Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de9f9c04-8478-4a35-b38b-bd27df643ce5 · outbound

This paper cites an unresolved cited work.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.369370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.369370Z digest=sha256:39fd70ebf0b7a6b8a12d1dc7732674367c14d00b19bd770bfc8c6a11fa637389

Observation 427d424f-591a-4f31-9921-131da82106cf · outbound

This paper cites Zico Kolter, and Matt Fredrikson.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Zico Kolter, and Matt Fredrikson

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:36.080670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:32.462686Z digest=sha256:fd4a6512a741e78af539fb9838cce59ed509a8e01192890e702c2c2adb68a798

Observation f90175f8-b759-464a-b357-e13422769c53 · outbound

This paper cites Dai, and Quoc V Le.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Dai, and Quoc V Le

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.576304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.576304Z digest=sha256:0522bb4faf0ea56e03ec612a92c9c26a16aa187bb32b52a0d041da889932d1e3

Observation 0df60811-3b14-484c-b131-d331a39ac21d · outbound

This paper cites Training language models to follow instructions with human feedback.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Training language models to follow instructions with human feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.728236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.728236Z digest=sha256:8b49dfb192fa61bb5fda1e89d65405b4bbd7cd2e2f5865f90c4657f6a887597f

Observation 4ef84feb-83b5-440b-b385-26ebd177d1e8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Direct preference optimization: Your language model is secretly a reward model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.818877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.818877Z digest=sha256:906b8db9a012061177aa47f41cf09e742fd0483e7b4d25c569eebf8f534cce23

Observation 2b992868-0ff3-44c0-b7b2-7794e294ae1a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.872161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.872161Z digest=sha256:fb99346daf4e7c064147f9ee2e3e946458ba66d69ab228c72af9d92abd68ed5b

Observation 5b6e2760-ce4e-47db-9af0-a46a4b894e1a · outbound

This paper cites an unresolved cited work.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:36.017467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:32.934431Z digest=sha256:9dd3882d83909b3e08643583a428d4b65e97d31368499443d33adadf1c8d698a

Observation c4aa1255-316b-49bd-9c0f-f8014a9a093d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Gemini: A Family of Highly Capable Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:32.985402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:32.985402Z digest=sha256:514cdad4b01c929935e09bc35dca0bbbae245d8c4723015e7a6980bdb7d73206

Observation 8713b269-1393-43b7-a7ca-55333114f870 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Lora: Low-rank adaptation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.040776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.040776Z digest=sha256:0c84a80aa5bf3b02b18f419d347496f014a59d4b97f1e3e0b6b226aaf0b26cba

Observation f634a9af-af2a-4454-bc05-8e043d5ce904 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.116622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.116622Z digest=sha256:edf7af18c6c5ba6bd86da4acad55d216888500b77c90c6804baffe4f719cba5b

Observation 960ad58d-a2de-4693-9111-0cb30b55e1e7 · outbound

This paper cites Pissa: Principal singular values and singular vectors adaptation of large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Pissa: Principal singular values and singular vectors adaptation of large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.171297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.171297Z digest=sha256:e78de4c39dffccc7ab485b23e89d6fd6564c574d23b485d292c09366f94ba0ec

Observation 7be6ac48-e3bd-41a8-bc43-8a79b5adbce5 · outbound

This paper cites Badam: A memory efficient full parameter optimization method for large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Badam: A memory efficient full parameter optimization method for large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.978076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.268272Z digest=sha256:50669fb386fdbc3abc1a597aa045270be4043a7156331d12c3c9dafa45b3d17e

Observation 7f30c869-55e7-4b8e-95a4-2b6cf3286f8c · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Visual adversarial examples jailbreak aligned large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.960806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.319627Z digest=sha256:c92ffe771afb1051f89170763540b52e45f448b2453d99ea18d95c11e3affef3

Observation 19e501af-73cf-4e17-9410-717948bfc92c · outbound

This paper cites Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.934984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.369132Z digest=sha256:8aaf8e1f824e102682f81c4493e2a267e558ec11a2c56b328b247a5fe45049a9

Observation 89dc442d-4a46-413a-bad7-eec925eab060 · outbound

This paper cites Jailbreaking leading safety-aligned LLMs with simple adaptive attacks.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Jailbreaking leading safety-aligned LLMs with simple adaptive attacks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.909348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.407286Z digest=sha256:dcfe1a2ee30c33ee83a1e0f22a832406a49e690f7c0cb0f0241cad4247c47117

Observation 9ff44d53-0ff6-4a37-a77c-cb02b38060c7 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.888574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.490650Z digest=sha256:54d53b68b212e382c94b05d81bce4911bb3de99ca5b6780d3260382b8d215498

Observation 070c56b1-b024-4089-b0af-2b10a0bbeb88 · outbound

This paper cites Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.866683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.549281Z digest=sha256:ee611ec4d97443360acf1ba0f4c0ebdce7d200bfe6ba012cfb0ea496bfd63a35

Observation 15a4efd1-f46c-4d61-8bf4-36c3999ba478 · outbound

This paper cites Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.841077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.627056Z digest=sha256:fb4a7739547becb520180a44c657eb2dfd1f6ada69d12be4fb0d6a2cfadaa05c

Observation 4950454c-2309-4f59-b35d-9bdb1fb441c3 · outbound

This paper cites Keeping LLMs aligned after fine-tuning: The crucial role of prompt templates.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Keeping LLMs aligned after fine-tuning: The crucial role of prompt templates

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.815897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.705062Z digest=sha256:0f0e2fc83eb23ec85c21507e002372d73bec457969a1e26664256cbb0144b84d

Observation 05d67407-e3c4-4ece-a8dd-92fa0b9dec88 · outbound

This paper cites Safe loRA: The silver lining of reducing safety risks when finetuning large language models.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safe loRA: The silver lining of reducing safety risks when finetuning large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.796442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.758826Z digest=sha256:7be2c0e52c62aac7c9e9d9136314b5d156900fb02a891c398ac7f4e1b870f4e2

Observation a6f6a29c-114a-4b38-8b5c-51670c8fb2c7 · outbound

This paper cites Tamper-resistant safeguards for open-weight LLMs.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Tamper-resistant safeguards for open-weight LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.778263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.806174Z digest=sha256:41c1e86f4edf3399c4b04ef05c941eb174718a8e757dfa8c850e51c077ca481d

Observation 66edb223-d40c-4319-8cb9-b637a74fb4f7 · outbound

This paper cites SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.759654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:33.917355Z digest=sha256:d6d7ad204ce5ee106d9fbd38a64ffac5a069e29690159d35b29c318e291553eb

Observation 532f6c18-5cea-401a-85e0-cf29774fab08 · outbound

This paper cites Safety layers in aligned large language models: The key to LLM security.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety layers in aligned large language models: The key to LLM security

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:33.998234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:33.998234Z digest=sha256:328b52d12e5182208a3e979b6471fcbd30efccde561b0e5f38f4419785db16cd

Observation 33abe46c-c251-4133-85d3-d9ab564bb628 · outbound

This paper cites SaloRA: Safety- alignment preserved low-rank adaptation.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SaloRA: Safety- alignment preserved low-rank adaptation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.730166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.037037Z digest=sha256:62834344b4ac7c179283e8d8b090d47ec3a08323e04ea1d4e13edd140f9b787d

Observation be3634a1-8495-4f10-b0bd-bc639ee0f10d · outbound

This paper cites Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.709450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.060960Z digest=sha256:f2030c090940b993938c64469fbe324481821ccc23475a85b17a7c8a933170a3

Observation 32b908e5-7d3c-4586-a750-fd4df42093f9 · outbound

This paper cites Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.692320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.090682Z digest=sha256:4dbf1c68e804f7ab4dc31042052557f7117302da40b78601438ab8c4d379a045

Observation dd8fa39e-d18d-4541-b2f9-f35aeb625878 · outbound

This paper cites Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.675398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.194340Z digest=sha256:e5558b89b1250605aa5bc3c18fde93bc75ab72f03b81a82a7648a7339c1338e6

Observation 5920d55a-850f-44df-8058-206b8b396850 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Safety alignment should be made more than just a few tokens deep

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.257702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.257702Z digest=sha256:2977ec783db0896911375f8a9cec64fe1ea79f95fce6359766631d43bf846e1f

Observation 063617ce-60fb-4340-ac7a-71aaf5b3622b · outbound

This paper cites Defending against unforeseen failure modes with latent adversarial training, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Defending against unforeseen failure modes with latent adversarial training, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.644201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.333322Z digest=sha256:b5987082f7e71b1420d2f05960eed870e7328c464676e9716081871ec24cc92d

Observation 479a383f-5ceb-4520-ae5f-c557d2eecf89 · outbound

This paper cites Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.627831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.410064Z digest=sha256:1255822ae04136f32a6e88f82094fb08dbc03312891da2ac1fa9e55582cca636

Observation 38cbbcc2-9445-4174-925c-36f6299ec164 · outbound

This paper cites Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation, 2025.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.601893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.478655Z digest=sha256:6ac46cba98f67ba77c9e9bee5f16dc9546b9e63b02d3230c559667b4c7059918

Observation 36b4a338-cdeb-46ba-8666-ca241a3f0cd2 · outbound

This paper cites Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.578472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.563306Z digest=sha256:56581687396f3e1ea01315b7845529e3bfd64ed021a549154298ae5d41a8b34a

Observation 0220af5e-d502-47e3-8036-525afa5ab5d0 · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.676929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.676929Z digest=sha256:a54740fd179ec854f2eb4c296ba42559d081893f12470802e76e9f8e5b1b92cb

Observation 98dead1d-25fa-4373-9a91-24834ff66963 · outbound

This paper cites an unresolved cited work.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:35.550215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.709566Z digest=sha256:b362dd5d9fbb0c294b95ab0741025f31c10013b4532ac0dab9565b2e17c3cc9b

Observation fc923ced-3d93-49ee-b673-1ffdbd439a66 · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifications.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Assessing the brittleness of safety alignment via pruning and low-rank modifications

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.524485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.714063Z digest=sha256:71b2ab8591589bbae818654989c2d01993f5f2f39befebd93337f43995591961

Observation 7f151a94-4360-40f2-b2d0-2976dc036c4b · outbound

This paper cites Toward secure tuning: Mitigating security risks from instruction fine-tuning, 2025.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Toward secure tuning: Mitigating security risks from instruction fine-tuning, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.505453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.718187Z digest=sha256:d7d006fa43bacee949f5849670d6e29a14383e74c069afe20b807db2a8f2be17

Observation 7f17451b-dd86-4a84-830c-b9a2252ea844 · outbound

This paper cites Locking down the finetuned llms safety, 2024.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Locking down the finetuned llms safety, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.486382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.721988Z digest=sha256:8e13bffc23c3b1498b4c9902135907ce5658a671f8981cfe111bd07ab5b03c7f

Observation c6d5ad28-eef6-4875-94b5-caa73e79424c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.726375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.726375Z digest=sha256:48a5f070d7ff8745d3d26f564ead84a040b687a72d8e58df6425b9373a135cd3

Observation 4e2e8df3-365b-4eaa-b24b-c22f448e7c63 · outbound

This paper cites Measuring massive multitask language understanding, 2021.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Measuring massive multitask language understanding, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.731783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.731783Z digest=sha256:132650c6ae2418cd63e0d75304577bc85b332400ad0b26983265f2454bc69b46

Observation 4edab1b0-6c34-4396-9b13-872213a50348 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Xing, Hao Zhang, Joseph E

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.738615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.738615Z digest=sha256:fb2a0cad16c66c7f148377c3b342ac12f1d7d8fcd32e1b18d0d32e27a59cc474

Observation 1bbfd0b5-32a2-4dff-95fd-f293d9622854 · outbound

This paper cites Certified defenses for data poisoning attacks.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Certified defenses for data poisoning attacks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.428950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.743466Z digest=sha256:273e99e803af97dc9291ec3ed89c2d92eac1f566b5f45e3166a070e21eb1fa3f

Observation 61c38025-b471-486b-b429-ff71b0240bfa · outbound

This paper cites Data poisoning attacks against federated learning systems.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Data poisoning attacks against federated learning systems

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.394671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.748082Z digest=sha256:249f5a40dfd083130fd32a1afa5c1a3037177f8cb5e748699bbed063eed34e2b

Observation c26f605a-8aa8-4e12-ad5c-382f47996d28 · outbound

This paper cites Immunization against harmful fine-tuning attacks.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Immunization against harmful fine-tuning attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.757469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.757469Z digest=sha256:e1206b92427b2d018dff9c23f0274d9a51a824b551e4d80c6c94ec5228970f48

Observation 77a852dd-4a36-4ce8-a082-12588c67ed17 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.761861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.761861Z digest=sha256:1ce2030551fa09a6766d2223c935bfd72ddfbfef03ac6a67cf3b0e93846bc8b7

Observation c7ec8314-dc2d-4267-8e4f-5a0ad634537f · outbound

This paper cites The case for an opinionated, theory-oriented real-time operating system.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning The case for an opinionated, theory-oriented real-time operating system

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.372479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.766914Z digest=sha256:0dd17abe3164d93fbba2b380f2782e93f0967ff85187da77728ce09fc333a7ef

Observation 0a34b9a9-a42a-4873-81c9-41aee11e8d93 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.771964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.771964Z digest=sha256:e1121713b9c575995748478315054568e179a6f486e911dae2332dd09d8c9572

Observation c324371c-2ac5-4ddb-b0ec-8e653cc155b7 · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Recursive deep models for semantic compositionality over a sentiment treebank

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.776353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.776353Z digest=sha256:626ed11be98c6419fe83ecd718b08ad222be2e15c42db3936385ec90b3b3a7a9

Observation 9ca2bf94-8217-4af6-8aa8-00a1a4362ebb · outbound

This paper cites Character-level convolutional networks for text classification.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Character-level convolutional networks for text classification

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.277370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.780136Z digest=sha256:64061e48be45ee85d839ca9ee1af37a0800596e76525f6d7a9784c5b31aa914e

Observation d4746240-e2a8-4e3f-947b-886a93bb3f2f · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.784487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.784487Z digest=sha256:c303e73b38e4823bce59678cab7ac7f5c5357d670d55992e2766af09715e8aee

Observation 5f838a6b-4c60-4c9b-90b8-d95bb9dde953 · outbound

This paper cites Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.789925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.789925Z digest=sha256:fed0368b897b64465aa0046d7e3df44b692fdbfac911d9565504bfb694f6db11

Observation 4890fa07-bb75-4ae3-8670-a364756f90c1 · outbound

This paper cites Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.795389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.795389Z digest=sha256:ba8c7ebc0a6d2ef74dc055197f6829d368b48eda24334b48234cc1673c832500

Observation b2893360-00ca-464f-b0f3-8d2aafcc35fe · outbound

This paper cites Alpaca: A strong, replicable instruction- following model.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Alpaca: A strong, replicable instruction- following model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.219475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.800016Z digest=sha256:a1c732497009798eab81819663b03b5171a62c5e90b17390d4356c2937b90641

Observation 463654fc-4125-455b-a554-7367748f69e6 · outbound

This paper cites Introducing the world’s first truly open instruction-tuned llm.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Introducing the world’s first truly open instruction-tuned llm

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.201334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.804830Z digest=sha256:ee35e0a92428622206b743ec3494396f23face09059378827a4a68ce3d5f9b7f

Observation 0e5921d9-feef-43d2-bd67-43c267879912 · outbound

This paper cites q_proj",.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning q_proj",

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.174645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.809340Z digest=sha256:12dfd48054ff97c6cdfac5b3c4875687064b28315e538ffee9a1c0b82255b061

Observation 498903de-128a-47f0-8494-6b248ff36c3f · outbound

This paper cites This is the most preferred response when confronted with policy-violating instructions.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning This is the most preferred response when confronted with policy-violating instructions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.815792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.815792Z digest=sha256:b39424b91da0ccdccb22e84872c28a3c1cf7fdf3a198d511a42fe660a3becce8

Observation feb12fda-29ea-44dc-a028-3c21b644ba91 · outbound

This paper cites It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.821541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.821541Z digest=sha256:7d7c40382f6b0d629e021e50a6d8831fa76d5b203fc3f7fa5b31c39f6f0777f7

Observation e23d4b7a-4399-4401-96c0-e9c832dcc3c1 · outbound

This paper cites It offers content related to the request but without embedding necessary precautions or disclaimers.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning It offers content related to the request but without embedding necessary precautions or disclaimers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.827143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.827143Z digest=sha256:80a9337ccb4b2a75981684e27bcdbbbc4f45031c1115670aeab467c06dbb395c

Observation 74b6152a-c0c0-4c9c-b8c3-eff51f401bda · outbound

This paper cites However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.832822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.832822Z digest=sha256:2c5d216d78c92acd6c24a4329344e0f3fb484444adf697eebe8f412140bf5895

Observation 0f610a2d-f3a6-44ba-99ce-8512d95179bc · outbound

This paper cites Candidate X Under Investigation for Embezzlement.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Candidate X Under Investigation for Embezzlement

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:35.072747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:02:34.839326Z digest=sha256:3a8c6e77826e7d7c6925ab888a2b020f677148383ca11de2ba74e7dad2bb90f9

Pith citing papers

No inbound Pith citation observations are available.