Pith. sign in

Paper Citation Record · LEDGER

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2602.06911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.06911 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:46:14.886511Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b3c9b1a-0ab6-4c15-a931-653918c625e9 · outbound

This paper cites (2024b) and He et al.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering (2024b) and He et al

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.881989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.881989Z digest=sha256:c04a92792ac105df6dfee05cf95607aec13239921e3595a7118eb74e8a6d75de

Observation fca2ddee-d6b6-4785-bc9d-149889d497c6 · outbound

This paper cites an unresolved cited work.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.886511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.886511Z digest=sha256:7c929d7bc3c8d74e7e60db708910b2b4f7ee9e268ae549c5fab31688df033b89

Observation cff9a125-00ba-468c-82d0-26bad94a0713 · outbound

This paper cites Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.790420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.790420Z digest=sha256:79e4c52d65e4d21565fe8591509e9de38846fade6e7300dd5973989823032b09

Observation 4568aedc-7d58-427e-81da-1c1653284209 · outbound

This paper cites Luxi He, Mengzhou Xia, and Peter Henderson.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Luxi He, Mengzhou Xia, and Peter Henderson

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.804852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.804852Z digest=sha256:d0d700ffc16fe2d0751fbcc031b8a808d6522645f2b8bf4acf74a3b8a92fd9aa

Observation 2abbcde9-9b3d-4e54-821c-3b459b676b2f · outbound

This paper cites URL https://openreview.net/forum?id= dp24p8i8Cg.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering URL https://openreview.net/forum?id= dp24p8i8Cg

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.809451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.809451Z digest=sha256:81c66ca98bb5a6a22265959d40d20397049d60d22f621f785e78d9582e4d6460

Observation 608c4094-fe7d-42db-91c4-eba22db57296 · outbound

This paper cites ISBN 9798400702310.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering ISBN 9798400702310

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.813546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.813546Z digest=sha256:3df76d03b591980d3cf5d08824c807a3bf40fcb4637b0edce787e8af2e2138b6

Observation 27f32768-a236-43d8-87c2-5e2116809be0 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.818615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.818615Z digest=sha256:82ebf336acbcefd4ee48ddcdfaf7c54ed5e66c97dd28627bc390c86f0af48cb8

Observation 2b2ab5f4-48a5-42c5-940d-b15044b12f3a · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.823403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.823403Z digest=sha256:3265a5ac1e088802fc9ce2179beb2922106cf295f8b9143f13ddfb18a9e00832

Observation cb19529a-d2d8-4a57-9bf5-95359a4f35ba · outbound

This paper cites Kyle O’Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Kyle O’Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.828547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.828547Z digest=sha256:78c7b6ae036d29bd6ccbc10d5ab29768d76aafffa968b22088ce2c27ba268f21

Observation cf606279-0b01-4581-b299-aef4377bcd57 · outbound

This paper cites an unresolved cited work.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.833451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.833451Z digest=sha256:c01698b4adf15e8001c74a232ee1254a651765002b676fdf2b37fba5cd79844f

Observation 608a7c66-ce8d-4b04-83d4-bfae2e8e242b · outbound

This paper cites GPT-4 Technical Report.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.838787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.838787Z digest=sha256:dfb2a684dfdf66dc794232fd875b6d0065e8174894516f6cb9b5939ef92e6a22

Observation 7013b613-6c98-4212-b40b-295c0e22d8b8 · outbound

This paper cites ISBN 979-8-89176-195-7.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering ISBN 979-8-89176-195-7

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.843583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.843583Z digest=sha256:5e241ee04e6f816c80f2d7f235ea95ec6fb2ea898b5aa42ab1e9ac31ce9d32b7

Observation 8193f5d1-1bd8-44cb-a112-509f52876962 · outbound

This paper cites Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.848303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.848303Z digest=sha256:db7df75b504e614ffa6072c838dcf5380333169ecd2b553be86c7919aa82297c

Observation 2fdc5205-36e5-4029-b521-a08d041a72e9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.852644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.852644Z digest=sha256:d6aa621a5afb0bf22549b90e922d94b4c54ef69300df614edcd6fdb6de749267

Observation 635d0d1f-45f5-4ec1-8ad7-6709e7139deb · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Jailbroken: How Does LLM Safety Training Fail?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.857511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.857511Z digest=sha256:338fa9ff7b363059ca4a59011407774cf0450c6cee3e9056e281b2a7620ffc32

Observation f31a21da-f524-4d98-b1b9-ad927f168a17 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Instruction-Following Evaluation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.862258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.862258Z digest=sha256:ce25e1629289b77f8721c4bee132317f99239c86e395873f0650bc25b71d1e11

Observation b65f4ac7-03a7-42dc-ae07-5bddd6793393 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Representation Engineering: A Top-Down Approach to AI Transparency

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.866706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.866706Z digest=sha256:0bb8fa654bd44f741c6ad3e37ea9e78a9e35858878f548a9828538a3ab470f1c

Observation f15d922c-090f-4bc9-b23a-810b20f1fdfb · outbound

This paper cites In contrast, MATH is only loosely aligned, reflecting its narrower domain and strict exact-match scoring.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering In contrast, MATH is only loosely aligned, reflecting its narrower domain and strict exact-match scoring

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.871822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.871822Z digest=sha256:857228ce10273a386c4c36421206d2b0b804026c201afe124d7ece19f7763a6b

Observation c6b62e94-65d5-4550-83c5-7ea3324b2509 · outbound

This paper cites Danny Halawi, Alexander Wei, Eric Wallace, Tony Wang, Nika Haghtalab, and Jacob Steinhardt.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Danny Halawi, Alexander Wei, Eric Wallace, Tony Wang, Nika Haghtalab, and Jacob Steinhardt

Reference 23

Resolution
verified exact
doi, observed 2026-08-03T03:49:47.372290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-03T03:46:14.800549Z digest=sha256:1a6426975e6976cb59099578827be101b2fde0f10286f3e7005b2bc5edcb2e39

Observation e417595e-ea0f-4b95-bf7e-fc2ac3fc5515 · outbound

This paper cites Program Synthesis with Large Language Models.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Program Synthesis with Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.775654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.775654Z digest=sha256:a36453b04d1527dd8f1621c2bc6c31cdaec0c4bf470eb6c1e05a9d23c15625bc

Observation b5f71dea-bc28-4e6c-b53a-e7ab0e8ac61e · outbound

This paper cites The current year is 2025.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering The current year is 2025

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.877674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.877674Z digest=sha256:42722996d8249b34592431ab31e93b6a1c508c5b1dc4f78c3d62788f67856db7

Observation 5c7b69ad-4a16-441c-9139-d0d94dad80b3 · outbound

This paper cites Eternal Sunshine of the Spotless Net: Selec- tive Forgetting in Deep Networks.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Eternal Sunshine of the Spotless Net: Selec- tive Forgetting in Deep Networks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.796043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.796043Z digest=sha256:04454d15adba9967f12f218bb76d056c3edf04c5dfd8b729e748add375f30fb9

Observation 79061286-4169-4c8f-8f42-9b3413502b80 · outbound

This paper cites URL https://openreview.net/forum?id=urjPCYZt0I.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering URL https://openreview.net/forum?id=urjPCYZt0I

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.785957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.785957Z digest=sha256:991ae866e489e6e06fa35245db45b13e363cf804ad398947d386b36a4afcb917

Observation e96d75aa-c237-4a1d-a88c-2607b78ab112 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.781234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.781234Z digest=sha256:32abb43ef954c325bc901b66fc0679da3757b5826365847bdbac123f9a9fbb76

Pith citing papers

No inbound Pith citation observations are available.