Pith. sign in

Paper Citation Record · LEDGER

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

As of 9 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 8 inbound Pith citation observations for arXiv:2502.05209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05209 v4

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.448064Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:06.683217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T10:46:17.055047Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf1f4201-9b52-47eb-a0b6-d553ec611530 · outbound

This paper cites write newline.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:14.990982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:14.990982Z digest=sha256:111f4c6110c7166c09f863ee277f651be10114f47c1354687f55d345c892ecde

Observation 183e7ffa-d049-4cfb-b8cf-0c3e179c6892 · outbound

This paper cites T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:14.997251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:14.997251Z digest=sha256:0c2519f092af256866e86e4ff390eae916f5af4794060765026a649536057343

Observation 73b1af65-f276-4185-8a02-8c85bf245625 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.002017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.002017Z digest=sha256:d7879c457f2b103d1308997a9d97e1602c9ce459257379fc59cbd60fbf5b6bff

Observation 30f4dc67-561b-4b2f-a9fa-dbd1744c6fb6 · outbound

This paper cites Many-shot jailbreaking.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Many-shot jailbreaking

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.007519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.007519Z digest=sha256:d503f136a6404bd8409e15a7f3da9740523ddb9a00acdcbbcc2e58f8cdf4fb7a

Observation c3e72b8b-62a8-481d-8271-b19d296bebaa · outbound

This paper cites Unlearning in large language models via activation projections, 2025.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning in large language models via activation projections, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.012554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.012554Z digest=sha256:f978dc3bc119f5c82b86ae27c588aac640e85ef25bbb8e375f3e0429d836a446

Observation b60765c1-c7ba-4511-b905-6f60d5639e29 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refusal in Language Models Is Mediated by a Single Direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.017432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.017432Z digest=sha256:5195a6c8fda2177ac9daa699028882e0739159fe1de0d9bdb3d6bf7d79691d3a

Observation 633bc806-8686-4c0c-bb0e-16e5805b4fdd · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.023992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.023992Z digest=sha256:e62b6e2297bceb5397d88b9bbb30faeae4ca2aa9db36743a3037a7175900c1a4

Observation c9d4dc73-3dd6-4f1b-be7c-ac44238e7d1c · outbound

This paper cites Open Problems in Machine Unlearning for AI Safety.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Machine Unlearning for AI Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.029821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.029821Z digest=sha256:cd16d340a88a8979a1302ea79ab1e7bdb79c184da6c14481dcd2f6fcc58e313c

Observation f6fcfcd8-651e-4314-9848-00fa399ccd0d · outbound

This paper cites Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.034861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.034861Z digest=sha256:b6971207d17c70dac192b7ef54b8b3321130f9a00d74f668c4f6913f6588531a

Observation e4f9795a-a64c-442d-95fd-3ba18fbca45d · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.040084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.040084Z digest=sha256:f3932a503af3ec0edea6f6373d88d32a7d34d70503baa2ac8fd2f6093dc13a0f

Observation a305ddd1-337a-4a51-a1d5-5426e98c3f80 · outbound

This paper cites AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.044779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.044779Z digest=sha256:02bc05f21727f11c2d06f217007c6b4407990e5b80d90e7b005250317963b9a2

Observation 9e2320bc-7954-45f0-8146-5a6b68311de4 · outbound

This paper cites A., Jagielski, M., Gao, I., Koh, P.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A., Jagielski, M., Gao, I., Koh, P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.049489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.049489Z digest=sha256:ea69b09c7a1d57479a3eca1e9f99337f943021b4862cc83bd08c21f76f768d51

Observation 375996cf-fb2f-4c96-8acc-b63e4b626094 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.054255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.054255Z digest=sha256:4a96f394f3eb8ae5124d8d0c73bf01ce6688e21771f9e5fcb7cf67a8dc299df4

Observation 9b042a6f-9ed9-45de-ada7-14c98f4f361c · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.059261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.059261Z digest=sha256:b8ed3c48269edd6aa0baa1c4989672af3c219e212111de61743d53fa4496e12c

Observation b5b04b81-e9bb-4b32-bb5f-815e9bd6bf24 · outbound

This paper cites Interim Measures for the Management of Generative Artificial Intelligence Services , 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Interim Measures for the Management of Generative Artificial Intelligence Services , 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.064636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.064636Z digest=sha256:705f2d7475e81aefb38a4be03fbed7f2739c0ce639fa91a5f225a9e85b89445c

Observation 037a7571-e26d-4a03-8e60-13c6c5568b60 · outbound

This paper cites G., Islam, M.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities G., Islam, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.069152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.069152Z digest=sha256:e2080517116111183607543a4bff29b5a6cce6e0ad9fa2ed3b3286d76e62fe85

Observation 9fd04bde-0107-41eb-b320-f5373135624a · outbound

This paper cites O., and Nilsson, F.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities O., and Nilsson, F

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.073705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.073705Z digest=sha256:d8939677b5ed40ed7fdb7b4b5ba23bce2bc8423c4258189b613e6f0e2f7de9b3

Observation 7f515d96-f150-4a66-87a0-e43671afe57a · outbound

This paper cites Do Unlearning Methods Remove Information from Language Model Weights?.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Do Unlearning Methods Remove Information from Language Model Weights?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.078108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.078108Z digest=sha256:53519c4394feb9a434d0b4fd27d5b51502db39c4c8ac912847ea8ef45975d4ee

Observation e625dda2-cebb-471c-a6ef-c43baa3c9637 · outbound

This paper cites The Llama 3 Herd of Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.082764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.082764Z digest=sha256:7805c7510b57338d86257558bbde99fb82f609361be792931c535125ad5fe6be

Observation fafdf18d-29b6-43e4-b8a6-8cf8436227fa · outbound

This paper cites The eu artificial intelligence act.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The eu artificial intelligence act

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.087548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.087548Z digest=sha256:2e91cbbda9a13024551872b9267a3c8f5849888a400fdd7016616559de073ac9

Observation feea78e1-b98f-41dc-94a2-5ebc51bc4061 · outbound

This paper cites Scaling Laws for Adversarial Attacks on Language Model Activations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Scaling Laws for Adversarial Attacks on Language Model Activations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.092289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.092289Z digest=sha256:e25d61108d4fc2e4839c2d8c0806f8c8dcbfdeda3b95f7a7355ea4704817b493

Observation 5b9d2a9f-2592-4aa0-9ac8-804cadcd21c5 · outbound

This paper cites Towards a science of ai evaluations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards a science of ai evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.097049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.097049Z digest=sha256:8a2e3f49973862385a919017c5172a15263475f968e18e786e9a2b2cbcf8d663

Observation db1812b8-3cba-45e1-ab28-f5d46d864bf3 · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Erasing Conceptual Knowledge from Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.101899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.101899Z digest=sha256:ac3a856b1bf43e58d5e2cc6ee5b995259f0000f5eab4e878af6513e4dd4c77d9

Observation ef386264-db9e-4fb4-bfb3-572786246e45 · outbound

This paper cites Stress-Testing Capability Elicitation With Password-Locked Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Stress-Testing Capability Elicitation With Password-Locked Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.106879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.106879Z digest=sha256:3eaafa2738fc26daf4b9ca66554a5d4621e8b86c62f68b0895c306a822f6f323

Observation ced76700-4f74-427c-98f4-f4fd392bcd53 · outbound

This paper cites Cascade: Exploring hierarchical inference in language models, 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Cascade: Exploring hierarchical inference in language models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.111767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.111767Z digest=sha256:958e69ff88e17879b64516b068156925fffc6a6bf8de3068fbce7dbba4c553f4

Observation 122f75cd-0c4e-4a4f-b5d9-096bb7b5eeee · outbound

This paper cites T., Haghtalab, N., and Steinhardt, J.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., Haghtalab, N., and Steinhardt, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.947064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.116500Z digest=sha256:539d64d412525ab954eb1ae28d5eecae0e3c352a13341d55e827508d3b8215b7

Observation 859381dd-9d09-4cc8-ab4a-ecce32ab1439 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.121434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.121434Z digest=sha256:a106a491086f8077633b0d11cfa7e5935af81d632461fd46c25a2c11b962f3ef

Observation a1b4633c-ab77-48bb-be91-52cd0af80adc · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.928848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.125861Z digest=sha256:4fadaeffb4f214c1c722731dc8b21d35b267ff20fd2a29ed132ad8a34d5036b0

Observation ac493279-4922-49e3-a797-b466f436aed1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.130612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.130612Z digest=sha256:0f9e63ec8440876f5f5c249a7217c3e7be7499c6d5065263c4dc4781f6666bd6

Observation c267204e-d9f5-4a06-a2ef-70836602af39 · outbound

This paper cites Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.135860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.135860Z digest=sha256:a2f3dff50ce3c791421f480e899c225fd0e65d2ff496ce2f00f5ae20175721fa

Observation 8dd47542-8160-46b8-a070-04a1f37247bd · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.141334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.141334Z digest=sha256:3d5e33d29cc36cecdf7625ee853b63c78479fb54abdacfac3952853c96396c9a

Observation 3bad5977-5ee9-49b1-a799-7ab9661c9872 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.146775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.146775Z digest=sha256:8dad1e6a399b7cf1f69ea6b76deea61f092664712c9ca888bb9c77b0d0a2f54b

Observation 853c3ab5-e289-40be-89ea-4a61d4329711 · outbound

This paper cites Language models resist alignment, 2024.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language models resist alignment, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.911752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.151928Z digest=sha256:9ac3f651dbf72d725dcb5215b06f93675346e5ce5b87f132e56b5c7158a7041e

Observation beaf5ec8-3b43-4564-b970-5b3fa8a64530 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.157036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.157036Z digest=sha256:a53ac651f79715867c06850a2d2ed0ec99f26d1c29a7109695efd6a74ad332a4

Observation 7c6eb7ce-c036-404c-b7b8-55c92326ee63 · outbound

This paper cites Act on the protection of personal information, 2025.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Act on the protection of personal information, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.894782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.161963Z digest=sha256:6857bf59e826ca126c24e6e74a972e3038c1fc0c5c3ad777ab15f403d635875f

Observation 8bb16e20-1d1b-44ef-82f2-e26f100e351b · outbound

This paper cites No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.166472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.166472Z digest=sha256:0a0d9cb15a41160f122b7c065a3171d6356c9f5b7501c65cde324f476acf6794

Observation 43acf8e1-359d-47e2-b21a-ca64edf4fd8d · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.171392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.171392Z digest=sha256:c062e260547c9222a16a45001bc1e23b81a2b607926d44722b91cbc26ba144e0

Observation e85b16c9-7c66-4b34-93ef-ddfa98ce1114 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.176057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.176057Z digest=sha256:ae8d52c8ac40c3fe7b14cefa6d3397dedba15674def33b84d6fbf2d060a26197

Observation c6c86131-280a-44a2-b768-ca80477cde25 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.181547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.181547Z digest=sha256:0c66884849f3727e568e17f5fb8cf1f759901fd57015f35179307f346b913e36

Observation 7569bbfc-337a-4e57-827e-c21e17128f0b · outbound

This paper cites Against The Achilles' Heel: A Survey on Red Teaming for Generative Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Against The Achilles' Heel: A Survey on Red Teaming for Generative Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.186818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.186818Z digest=sha256:aef12b360923c5e7b22d990bb99c22f43ce54be8d357602bf53fbdd62a420641

Observation c0c5d3e0-6632-47ea-a285-1a98c0d9d47c · outbound

This paper cites Continual learning and private unlearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Continual learning and private unlearning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.878763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.191766Z digest=sha256:40b7ac3f9c7ae5db94c50ecdc7f738b3cd06b8f0943700e1e7f8503f89bc5c86

Observation 9a67d7ca-f77f-4652-8175-fc201462abfa · outbound

This paper cites Rethinking Machine Unlearning for Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Rethinking Machine Unlearning for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.196404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.196404Z digest=sha256:20fddfac6d4e39035cd2c4a67c92f62cbbb45ac159eebda90f896f6491eef706

Observation 153ff993-0b35-4b9e-a6a0-c6be10f7c39e · outbound

This paper cites Threats, Attacks, and Defenses in Machine Unlearning: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Threats, Attacks, and Defenses in Machine Unlearning: A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.201431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.201431Z digest=sha256:aed1ef76a65601029615c81315b6e5761a2593b82cc62459b743df6b1180a681

Observation dfa340c9-2760-493a-9f3e-6fed376a1db0 · outbound

This paper cites Large Language Models Relearn Removed Concepts.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Large Language Models Relearn Removed Concepts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.206547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.206547Z digest=sha256:c1d8a6e02bb16d05894f2918b2b4852e1da751975eccda10dc2c67e79a074a8a

Observation b1a68a4a-4299-4fca-ba6e-6f07c0521583 · outbound

This paper cites and Rimsky, N.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities and Rimsky, N

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.862145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.211427Z digest=sha256:8a4f0e57421b73d56a6befc3ca7ea741aee2ac4704244413af3d13da249e80f8

Observation 52289a63-379a-44f5-b904-0db3406e363f · outbound

This paper cites An Adversarial Perspective on Machine Unlearning for AI Safety.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities An Adversarial Perspective on Machine Unlearning for AI Safety

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.215964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.215964Z digest=sha256:599e72d6dab77ed628f6d6131438e9ab4e2b6b60182f7d2df1a13bcd35b628fc

Observation c24c5f8c-9705-45b0-b204-b9a06d97a898 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.221048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.221048Z digest=sha256:d1de82441bed93322351b26bea5794e7df79a385dfb9cca0bf4e5a03dbbae094

Observation 4ce66b5b-93f5-4eb5-883c-4f0784adca3a · outbound

This paper cites Pointer sentinel mixture models, 2016.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Pointer sentinel mixture models, 2016

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.225782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.225782Z digest=sha256:6aa19fe6fd62f0b7c4f3c0dbf9a63d47bf7b29ec94dd80b891ec7ec2a9787168

Observation 8a9b5cb3-561d-4abb-83f5-92739136e69a · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.833753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.230458Z digest=sha256:bf38eb87f61ca9dbb17b0b856df2f5959165830c1d76ac7472a3b0f1bc696336

Observation f24989a3-4478-4e69-aef5-a8658535b9d4 · outbound

This paper cites AI Risk Management Framework : AI RMF (1.0), January 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Risk Management Framework : AI RMF (1.0), January 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.817088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.234922Z digest=sha256:2ac39ba758f4ae45dffee59f2c5bfbf041808cb2007799f0ee34bc653af47d65

Observation 94250e00-55ce-41fe-82d0-0e93f89049fe · outbound

This paper cites Openai system card: December 2024, 2024.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Openai system card: December 2024, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.799205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.239688Z digest=sha256:4dda0a556d4a27a321cad0bb16cb99019d847cf3a1d2b874582457c3dcdf343d

Observation 11943eb2-7684-42a4-9402-04f4fe61a795 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.244674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.244674Z digest=sha256:79873128a903ad8a2498f24f62610068595d609762d53ce510726092431b7ee1

Observation 45cc7030-3d76-4a04-a0b8-f5b5f5e44703 · outbound

This paper cites Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.249627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.249627Z digest=sha256:127f27122e09ce1cb6cfc62d30d6ba4bddd7e23bee2eab7d74d9b6e043416801

Observation a9921514-bedd-4a3d-a40f-c3ecc1d08f61 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.254462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.254462Z digest=sha256:4ef51dd5bb538266227a784079ef58aeaf45759d92d4311b3d68251755bd7050

Observation 7657da45-83b4-4452-96cc-aaf6b36efad7 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep, 2024 a.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Safety alignment should be made more than just a few tokens deep, 2024 a

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.783616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.259603Z digest=sha256:5ece6482b7ef6be2c14161a819f3f38f0edd94561df3e4b5a3db3ea9ae5eb2ee

Observation 46082899-8732-4da9-ad97-7c6104d9b44c · outbound

This paper cites On Evaluating the Durability of Safeguards for Open-Weight LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On Evaluating the Durability of Safeguards for Open-Weight LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.264207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.264207Z digest=sha256:84730c7bbe6dcb47b96403bb12efdc543f978a9f740b561e6558cbe418d85cdf

Observation 46e855c5-3c66-40d2-b412-a59fd126938b · outbound

This paper cites D., Xu, P., Honigsberg, C., and Ho, D.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities D., Xu, P., Honigsberg, C., and Ho, D

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.767226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.269083Z digest=sha256:dad9ff7edfec89cc02554450aaeb86dc166d1be41c4d3d5ca8d4f96861b78535

Observation 02e2e5a7-928d-45c4-b7d0-34193dd1bf67 · outbound

This paper cites Open Problems in Technical AI Governance.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Technical AI Governance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.273714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.273714Z digest=sha256:12a80a067bf329488ac5eb165d97ec2802affb5114d4c94b2e36128b7571f0a8

Observation d1c44e37-5f68-4494-8618-560ad26d944d · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.278636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.278636Z digest=sha256:0b44638b2b0c8d38896c39605a1d6666da0a930502d5fd4ac76d45132536677a

Observation dc0996cd-9a84-41e8-95b7-8e63d29f3799 · outbound

This paper cites Fast Adversarial Attacks on Language Models In One GPU Minute.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fast Adversarial Attacks on Language Models In One GPU Minute

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.283768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.283768Z digest=sha256:7ecbe3a10a6a36df2285e6b1639e6b4666a760b5c1319cb1f3c576e9c83c8d41

Observation 494f9dc6-20ee-48f4-825c-e56e317b936b · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.749567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.288872Z digest=sha256:686046636c7f4dfd84585d842b239c22f7f16f6b8e57e5f44cd2d5748a083568

Observation 77db4e86-0d51-49ef-b89a-be27da8b1815 · outbound

This paper cites Towards best practices in AGI safety and governance: A survey of expert opinion.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards best practices in AGI safety and governance: A survey of expert opinion

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.293676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.293676Z digest=sha256:cb4cfc6ac7814a883e4023bfba4f20267a674cab7913bef6c852caf712d56502

Observation 6ccc3f80-0418-4903-a57f-661b4c4fc24c · outbound

This paper cites Adversarial attacks and defenses in large language models: Old and new threats.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial attacks and defenses in large language models: Old and new threats

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.730884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.299109Z digest=sha256:a96ae99f85bc6063b94e5227581298913885f5817331a17d04cc54bb544f9b29

Observation ad40a32b-3f19-4bfc-b54c-e45792a8415a · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.304017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.304017Z digest=sha256:9f52a5a7d2c2b453b378d27e500aaef7334d2a2d338ac60029941ad1f431d494

Observation af0cb92b-f7bd-42ad-9c34-e3a1ece7f36f · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.309285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.309285Z digest=sha256:bd476a9456acdf4ed7efe8256f56e78b33ae85541d7becff7815c98cea9b084e

Observation 5b9b18c3-6020-455f-9cdc-8ab91d420ad9 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.314358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.314358Z digest=sha256:5fd0f7e2863f103d172b3a69d1c50c4b5a6a44b33a6868e7d30aeaa382cfab08

Observation e882aafa-ceab-4c62-9b51-c96884513ba0 · outbound

This paper cites Model evaluation for extreme risks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Model evaluation for extreme risks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.319250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.319250Z digest=sha256:10742d6bddc0868768c917cad3d42aec27a21edc26702a66d482b54180187259

Observation 04b13167-b951-403d-9833-4c2ac64427dc · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.324252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.324252Z digest=sha256:ca3df61c600caa297cf20b3429745f0ec8a0239e9b75074766197a959ae978f8

Observation 8b1f053a-be81-4969-aa6a-d73c134f95ac · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.329073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.329073Z digest=sha256:df7dae4fe59618f2b6693ae3d4bf16246adb0ce6c84f5ba5e6463c5b8b50edb6

Observation f6df5fa0-756f-4281-87e5-1492cc8cc155 · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.714557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.334078Z digest=sha256:af87960664e8b48fdef8bd42a8968b116238be13ccad4581c5ef40e31eadeeab

Observation 3b01b132-ce61-46a2-aedf-4dcea52fdc81 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A StrongREJECT for Empty Jailbreaks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.338954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.338954Z digest=sha256:9b493ffe45f799799bc826aef6264faa5e734b1970b286aad3f2bb3afbb0b6ef

Observation 1d3ddeba-b616-47bb-bed5-613a2bf254ed · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Simple and Effective Pruning Approach for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.343866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.343866Z digest=sha256:077818b452c1449b21bd32218df6644e1af7c5a9099f5ce12badc428a6f5b7a6

Observation 9a2486f0-3e1b-4f9f-b395-d26a1835bd4d · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.348837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.348837Z digest=sha256:82e1fe170e48dee1b2303ef8a3fd1f71b2fab8cdf1119d2a3f815d3bb39c1167

Observation aa727472-dfef-42ec-84b1-83bb5e94f6d2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.353590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.353590Z digest=sha256:507dba7fe72127c620547cf8849a51f26656fbf6414612981a8de0e8cccf8a5f

Observation 1dfc3555-196d-425e-833a-f95e4ee86445 · outbound

This paper cites A pro-innovation approach to AI regulation.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A pro-innovation approach to AI regulation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.698068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.358517Z digest=sha256:2a104a998b013b7ea6b3a31a39017e76a06e084d6777c903936d3be19af9b2a7

Observation 1876fddd-203d-463e-839f-4af8685db792 · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.363084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.363084Z digest=sha256:403592264731676393cc728d74aee9ee8689c07220ea6b5b87a30cb7efeab655

Observation dd41a068-d301-48c4-a177-4c4270a48e03 · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.367750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.367750Z digest=sha256:a14459c3b3fb3d41451f2c5855a21c02e9f28c607d558a66ff6e9ce8e673a500

Observation 0bf8c76c-ecb1-4d01-a8c8-a7a07a00873c · outbound

This paper cites Efficient Adversarial Training in LLMs with Continuous Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Efficient Adversarial Training in LLMs with Continuous Attacks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.372451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.372451Z digest=sha256:4d5072a4d06767c59a5dc50209c6b244dc9f9309aec4594ececa2d15475f8243

Observation da6e0ecc-79bb-47cc-80ab-863b8edc129c · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.377658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.377658Z digest=sha256:9b58629ef9acd308484b012fecae3a59fb59af6390276d7c1e3bfbcac8a51751

Observation fc9df143-d719-4af9-8b61-8e95f72f1ce9 · outbound

This paper cites On the vulnerability of safety alignment in open-access llms.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On the vulnerability of safety alignment in open-access llms

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.680996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.382845Z digest=sha256:36173e01dd40ce2c20201b40d1a945cdc5b67d3d734949cde912c03370775b6f

Observation c0b52edb-205f-49c1-879f-be6937fc0b5e · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.387716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.387716Z digest=sha256:6fa824917cbe40ca4753e9f1abf6227942e09e1feb0a9d8c3b70f9e1f5af03cd

Observation bc50e58e-37ee-492d-bc8c-52347c806179 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Low-Resource Languages Jailbreak GPT-4

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.392539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.392539Z digest=sha256:428e670b87b2d120657a7eaabca71af528a53f4d7df0a5759b6d5c02ac8faca9

Observation aefc68dd-f5a7-4218-8fe3-77f7c6f34395 · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.397309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.397309Z digest=sha256:6d91927fddb31d5750127b50a136388e31996f2035b2a197934189bd746d8c9e

Observation ad4eb8ad-9be3-49ed-bbd1-b106c3b322d5 · outbound

This paper cites BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.402376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.402376Z digest=sha256:b4351fd4a5988918555d2209d56fb475d3e150d8d43f0c423f51734cfb022eff

Observation 75cfd0a4-16c1-4066-8ef6-5fb86992c319 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.407456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.407456Z digest=sha256:b61938217cb05db8987dfe2990ecbae742f5904fa5e64d241af4b7b05bbec027

Observation f3323a75-f9ad-4ce3-b497-b1bad5992b69 · outbound

This paper cites Adversarial machine learning in latent representations of neural networks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial machine learning in latent representations of neural networks

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-08-09T14:47:24.387209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.412546Z digest=sha256:67f2fd1ce0021baf151b25629e0fc5cc3d49b64b989d4119e5f3caf13978cdb7

Observation 73905ea9-a758-443f-8657-588ec377ee6c · outbound

This paper cites Catastrophic Failure of LLM Unlearning via Quantization.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Catastrophic Failure of LLM Unlearning via Quantization

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.422836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.422836Z digest=sha256:0ee27e1eba1e3afda85c5fb66d0b4665559d9dc0e6f26acb13fc6593217ab152

Observation c52482b9-2f28-4048-80f8-85433775e0de · outbound

This paper cites A Survey of Recent Backdoor Attacks and Defenses in Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Survey of Recent Backdoor Attacks and Defenses in Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.427626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.427626Z digest=sha256:bf764e984ff905f3096f5d017cf77f4dbc2c40e26a5d0644242474a3dd77e65a

Observation 0f86f630-1d5a-440f-a596-014d8a74a2cc · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.432963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.432963Z digest=sha256:3c2332cee38a19c35568898e8e91f787d976230f85488b8b158d26c195290aef

Observation d94a0715-f578-490c-8a47-ec5e68464ee1 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Engineering: A Top-Down Approach to AI Transparency

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.438083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.438083Z digest=sha256:1764c89da904e45b4a4058b30914fd292a2b535fc2deac7389c2df71078f09d4

Observation c0a5011b-8beb-48f3-b3fd-c6ba8f62a5b6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.442928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.442928Z digest=sha256:e6f298faf1fa032ca131fcaf1abbbc2b3a1341e92ea04eefdaa0b32ec960c0f8

Observation 73f8f072-4083-4492-8bce-0cafd19c6284 · outbound

This paper cites Improving alignment and robustness with circuit breakers.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Improving alignment and robustness with circuit breakers

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.664850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.448064Z digest=sha256:ad7345ecfc6764a6816ea358070b4930c4520cd2ec56aa2d9f0b50487e0ed003

Pith citing papers

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · inbound

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods cites this paper.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:9b985ca14abddd2f2715cc2aedec27cd86c93288b133b614e6f194b7785f89f0

Observation 5b691c14-d0ac-4408-84b3-1a983d3c9287 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.212266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.212266Z digest=sha256:50c4f640f6b3cc94c01fb6ae5ce8c176047268485af42bee67120487275afe3c

Observation 0abc5024-ef5d-4655-95da-9c9b68ade291 · inbound

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning cites this paper.

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:46:17.058220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T10:44:53.516653Z digest=sha256:c08b8c71049e609fb8f636b52e593b852d13395a8a7b31882d35306b8d608306

Observation 9195ab00-e824-41a4-9ef9-dd9eb8448c61 · inbound

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories cites this paper.

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:41:33.209039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:41:33.209039Z digest=sha256:9afa60b7b47f6464ce1f6c420a7d3823ce7a0194a0b0b1175cc5035e53914e1d

Observation 5ba3b8fb-a6c3-43a7-8dbe-1d7aa8abb55a · inbound

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns cites this paper.

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:59:01.358929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:59:01.358929Z digest=sha256:7e768812d58e9bee6fd4e970277f2cc2d32ebaa872503246ac0fcf82c4a14008

Observation 735466bb-16e9-47f0-ad25-d74394d5c890 · inbound

Operationalising the Superficial Alignment Hypothesis via Task Complexity cites this paper.

Operationalising the Superficial Alignment Hypothesis via Task Complexity Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:49:13.999922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:49:13.999922Z digest=sha256:d6abcff55bf89c26bb8b470e3a74b17345364dc641782a7c7f1134f05861da80

Observation 9f63ce60-b49b-432b-94be-3091cd96bea2 · inbound

An Independent Safety Evaluation of Kimi K2.5 cites this paper.

An Independent Safety Evaluation of Kimi K2.5 Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:43:11.577719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:38:18.674355Z digest=sha256:91b8014d7d06479d8eea0bea5554d7b3e035f601dd6355be2532e9426695b290

Observation ec539424-962b-4ec3-b385-e3a4912b5b1d · inbound

Is your algorithm unlearning or untraining? cites this paper.

Is your algorithm unlearning or untraining? Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.104808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:06:08.962042Z digest=sha256:adfd070da63d96259be68ee2b58960fb5a6db90802ae3be757ceb7b6b0088051