Pith. sign in

Paper Citation Record · LEDGER

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

As of 10 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 8 inbound Pith citation observations for arXiv:2502.05209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05209 v4

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.448064Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:06.683217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T10:46:17.055047Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf1f4201-9b52-47eb-a0b6-d553ec611530 · outbound

This paper cites write newline.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:14.990982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:14.990982Z digest=sha256:a1322d737dfa7072b194e2603cdf269c12a7c1fcf51147503b22b4150f8a79a4

Observation 183e7ffa-d049-4cfb-b8cf-0c3e179c6892 · outbound

This paper cites T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:14.997251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:14.997251Z digest=sha256:75aa4e34a2232bd980c34491769b6579fe39aa3927e89c73cbe135c61197b435

Observation 73b1af65-f276-4185-8a02-8c85bf245625 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.002017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.002017Z digest=sha256:ecb7cbc216f712132821b066747b5e8c8d32742089de99ae51a40f7c24788538

Observation 30f4dc67-561b-4b2f-a9fa-dbd1744c6fb6 · outbound

This paper cites Many-shot jailbreaking.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Many-shot jailbreaking

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.007519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.007519Z digest=sha256:5e947eace493200792d596a62c1a6553ab165d69599a5ce7d9c16832d164705a

Observation c3e72b8b-62a8-481d-8271-b19d296bebaa · outbound

This paper cites Unlearning in large language models via activation projections, 2025.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning in large language models via activation projections, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.012554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.012554Z digest=sha256:266a034b40deae7436ffb120dd8b80c5081716540c063ab2be10aa61ff466378

Observation b60765c1-c7ba-4511-b905-6f60d5639e29 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refusal in Language Models Is Mediated by a Single Direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.017432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.017432Z digest=sha256:e5b2d49773be8e4a3df639b9bd53c34dbd3df486bcaec9ca920deb244ef7cca9

Observation 633bc806-8686-4c0c-bb0e-16e5805b4fdd · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.023992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.023992Z digest=sha256:8c9e92c735c4eca5843395afec937a0475f05dcb5cafa9a72e53a3473480a398

Observation c9d4dc73-3dd6-4f1b-be7c-ac44238e7d1c · outbound

This paper cites Open Problems in Machine Unlearning for AI Safety.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Machine Unlearning for AI Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.029821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.029821Z digest=sha256:85efd51fd0f1129a6c7afd29a892b625e66b994e92b21b5ead78d255a7bd2cf5

Observation f6fcfcd8-651e-4314-9848-00fa399ccd0d · outbound

This paper cites Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.034861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.034861Z digest=sha256:f0caaddf2cab56010fc396ed868e8b47ee8ad689bd7e056a16820b21a4dd32cf

Observation e4f9795a-a64c-442d-95fd-3ba18fbca45d · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.040084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.040084Z digest=sha256:239a73b81ea730de8234bd01a4ef3acac3439f2301ed328225cc918be247e1dc

Observation a305ddd1-337a-4a51-a1d5-5426e98c3f80 · outbound

This paper cites AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.044779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.044779Z digest=sha256:8465ab7d748a4901b25dd47c334f1985f3746821c9a105596d80e0fa2ddaf1ac

Observation 9e2320bc-7954-45f0-8146-5a6b68311de4 · outbound

This paper cites A., Jagielski, M., Gao, I., Koh, P.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A., Jagielski, M., Gao, I., Koh, P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.049489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.049489Z digest=sha256:b676484b53159dc9be3eeb402073e17b3f44bcc924558dc9de79196d2680ecc4

Observation 375996cf-fb2f-4c96-8acc-b63e4b626094 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.054255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.054255Z digest=sha256:b9abf0a786013136579f87ea6f35406a43fb8ae77662df4c59b5b0818d6622bf

Observation 9b042a6f-9ed9-45de-ada7-14c98f4f361c · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.059261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.059261Z digest=sha256:821a8c18e534b9e1d7355f386f6e14252cfc498e110c31fb6e8b94c370ff6e59

Observation b5b04b81-e9bb-4b32-bb5f-815e9bd6bf24 · outbound

This paper cites Interim Measures for the Management of Generative Artificial Intelligence Services , 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Interim Measures for the Management of Generative Artificial Intelligence Services , 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.064636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.064636Z digest=sha256:9ca3990f59d0715886d5aef1e35afd442ef8083cdb0962380713e738ca0376ba

Observation 037a7571-e26d-4a03-8e60-13c6c5568b60 · outbound

This paper cites G., Islam, M.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities G., Islam, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.069152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.069152Z digest=sha256:c886500eb6d04a04d6818425206974024c1a38879fdad3efc762d0c66cfe05a5

Observation 9fd04bde-0107-41eb-b320-f5373135624a · outbound

This paper cites O., and Nilsson, F.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities O., and Nilsson, F

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.073705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.073705Z digest=sha256:36006440633454460be07842e59c62bed74c32b8c0172f363a36d72d98629230

Observation 7f515d96-f150-4a66-87a0-e43671afe57a · outbound

This paper cites Do Unlearning Methods Remove Information from Language Model Weights?.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Do Unlearning Methods Remove Information from Language Model Weights?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.078108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.078108Z digest=sha256:27dc493d5fd737802d6209e5536ca035d654aac3e21c705a6db6aa0f9f1fc094

Observation e625dda2-cebb-471c-a6ef-c43baa3c9637 · outbound

This paper cites The Llama 3 Herd of Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.082764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.082764Z digest=sha256:97852e28cc01fb30e378020d8d1b1a3d0cd48f889554672c84a1e3df0599f816

Observation fafdf18d-29b6-43e4-b8a6-8cf8436227fa · outbound

This paper cites The eu artificial intelligence act.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The eu artificial intelligence act

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.087548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.087548Z digest=sha256:f7bd6c812d54a24905706b42c60bd94b8d735760af21035e501b6da1bac1277d

Observation feea78e1-b98f-41dc-94a2-5ebc51bc4061 · outbound

This paper cites Scaling Laws for Adversarial Attacks on Language Model Activations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Scaling Laws for Adversarial Attacks on Language Model Activations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.092289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.092289Z digest=sha256:a1fdbb2ffeb9f7fff6c38d89f970c138e8078e2f09ab86b9d6393ca8f3e7dd31

Observation 5b9d2a9f-2592-4aa0-9ac8-804cadcd21c5 · outbound

This paper cites Towards a science of ai evaluations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards a science of ai evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.097049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.097049Z digest=sha256:0338d834dace1ca7851c6630bd064ff40b8819dd3383c01de8944904bb96ca0c

Observation db1812b8-3cba-45e1-ab28-f5d46d864bf3 · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Erasing Conceptual Knowledge from Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.101899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.101899Z digest=sha256:fa62716d8a2cb66b2b299e117340c6a1e22ecb52466c6e2a55ef8948aeee5fdb

Observation ef386264-db9e-4fb4-bfb3-572786246e45 · outbound

This paper cites Stress-Testing Capability Elicitation With Password-Locked Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Stress-Testing Capability Elicitation With Password-Locked Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.106879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.106879Z digest=sha256:26c48596a6dae92a3b09ac81d4e01dbc45a6b66de01e2d9fcf381db406d46981

Observation ced76700-4f74-427c-98f4-f4fd392bcd53 · outbound

This paper cites Cascade: Exploring hierarchical inference in language models, 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Cascade: Exploring hierarchical inference in language models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.111767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.111767Z digest=sha256:98fa2d5997933c8c6dca1aefeecd662f654f171f440168302b0594b23619a2e7

Observation 122f75cd-0c4e-4a4f-b5d9-096bb7b5eeee · outbound

This paper cites T., Haghtalab, N., and Steinhardt, J.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., Haghtalab, N., and Steinhardt, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.947064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.116500Z digest=sha256:90632b45f03aa22b72854dff77f585d97e0e3e478ef9dae24d85b4c442a37f43

Observation 859381dd-9d09-4cc8-ab4a-ecce32ab1439 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.121434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.121434Z digest=sha256:b5db97fd2dfd0251403ba83099edd3a80e43db14e3fc322ec965150eff07f837

Observation a1b4633c-ab77-48bb-be91-52cd0af80adc · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.928848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.125861Z digest=sha256:fd0104ad57b3ed39fd27cae4d8c64923a561fcad039cf3aa102e60a3a27b1f1d

Observation ac493279-4922-49e3-a797-b466f436aed1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.130612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.130612Z digest=sha256:c9c778379a48289f6aad3c7059d6f4ee9ddbb619708ec6175e5ef612826ee98b

Observation c267204e-d9f5-4a06-a2ef-70836602af39 · outbound

This paper cites Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.135860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.135860Z digest=sha256:54e6ff6cf3fa47144ea5d5e36c8dd16060de5865dab02d824f71848470c82323

Observation 8dd47542-8160-46b8-a070-04a1f37247bd · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.141334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.141334Z digest=sha256:9c2c920099b97e645573b518dc29226221e307881200bd0f62c22ab79518e934

Observation 3bad5977-5ee9-49b1-a799-7ab9661c9872 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.146775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.146775Z digest=sha256:34ab65a453b397093942a138eaf68c5c580369a8e67d89ca77ef67c3131d266a

Observation 853c3ab5-e289-40be-89ea-4a61d4329711 · outbound

This paper cites Language models resist alignment, 2024.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language models resist alignment, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.911752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.151928Z digest=sha256:f0b320a1b0a0471483d61bef0134146d240179ce7123c2e3a5c5cd74a3f1f74b

Observation beaf5ec8-3b43-4564-b970-5b3fa8a64530 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.157036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.157036Z digest=sha256:9b351a5b9da098a25613067b896d79b65052c17894a4f3fab2b8a917ccaf423a

Observation 7c6eb7ce-c036-404c-b7b8-55c92326ee63 · outbound

This paper cites Act on the protection of personal information, 2025.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Act on the protection of personal information, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.894782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.161963Z digest=sha256:7d0dd3449d8d8ed5695a48ec8e537e09dfaed8a0ec7d0125e22439a87260f8e6

Observation 8bb16e20-1d1b-44ef-82f2-e26f100e351b · outbound

This paper cites No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.166472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.166472Z digest=sha256:c7aa3237f96cad5f12520fbe4b3747233f07eb577613dde48f52c286185ccc3c

Observation 43acf8e1-359d-47e2-b21a-ca64edf4fd8d · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.171392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.171392Z digest=sha256:7a80ad9e66e8478081f6e13c1d14cbc3a252c68cd9f20e1858a9af0fcb50990c

Observation e85b16c9-7c66-4b34-93ef-ddfa98ce1114 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.176057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.176057Z digest=sha256:32d5843dd4021c86c3e77a06363dd9b4525ff58c0367c721c8f230e773558d3f

Observation c6c86131-280a-44a2-b768-ca80477cde25 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.181547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.181547Z digest=sha256:bc0d1e709d408873d0a6ab8fb1e994d59e125f3115c229dba9a6c75e08474adb

Observation 7569bbfc-337a-4e57-827e-c21e17128f0b · outbound

This paper cites Against The Achilles' Heel: A Survey on Red Teaming for Generative Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Against The Achilles' Heel: A Survey on Red Teaming for Generative Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.186818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.186818Z digest=sha256:6966b5840eed0ed1a28681896fce7c1043c27b77cdec3feac744dfabf016ea39

Observation c0c5d3e0-6632-47ea-a285-1a98c0d9d47c · outbound

This paper cites Continual learning and private unlearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Continual learning and private unlearning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.878763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.191766Z digest=sha256:61923adb732c1084b63b9b66b17e64ccfedec80f9f32d15cd3b33c7976c7f6cd

Observation 9a67d7ca-f77f-4652-8175-fc201462abfa · outbound

This paper cites Rethinking Machine Unlearning for Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Rethinking Machine Unlearning for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.196404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.196404Z digest=sha256:feee2036d4884821a5f76a511ec4d2ae608f964dba6cdfbb1efcb610d88ab6fd

Observation 153ff993-0b35-4b9e-a6a0-c6be10f7c39e · outbound

This paper cites Threats, Attacks, and Defenses in Machine Unlearning: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Threats, Attacks, and Defenses in Machine Unlearning: A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.201431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.201431Z digest=sha256:cd8b1c0c171e4cc78162857b1886b943e295077b3a3461862c30722e1be5ae33

Observation dfa340c9-2760-493a-9f3e-6fed376a1db0 · outbound

This paper cites Large Language Models Relearn Removed Concepts.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Large Language Models Relearn Removed Concepts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.206547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.206547Z digest=sha256:a1000d0cc192326aa8eae20651183a9405374565a850b656425084f28f99d9d3

Observation b1a68a4a-4299-4fca-ba6e-6f07c0521583 · outbound

This paper cites and Rimsky, N.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities and Rimsky, N

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.862145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.211427Z digest=sha256:887f1336d68c2809b3ea0110554e3fd57798a63593f5c5411d192875f1072381

Observation 52289a63-379a-44f5-b904-0db3406e363f · outbound

This paper cites An Adversarial Perspective on Machine Unlearning for AI Safety.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities An Adversarial Perspective on Machine Unlearning for AI Safety

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.215964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.215964Z digest=sha256:4b4fbcbc9bfb2a37259b508855ee75342d56d495d6fc493f032314b57c243f52

Observation c24c5f8c-9705-45b0-b204-b9a06d97a898 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.221048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.221048Z digest=sha256:044b2571d99470fb28e590a5638585452deae96d0f38a55c3093f65ae66118da

Observation 4ce66b5b-93f5-4eb5-883c-4f0784adca3a · outbound

This paper cites Pointer sentinel mixture models, 2016.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Pointer sentinel mixture models, 2016

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.225782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.225782Z digest=sha256:825ec3208c69f2eeb66b26cb0cdea39bd61b251abcde30bc7f6120afbdac4d58

Observation 8a9b5cb3-561d-4abb-83f5-92739136e69a · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.833753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.230458Z digest=sha256:ee62c717a9f7c75ea883d26e3210c8dd0f99a07c37b8f0827bf5e9c1de9d4280

Observation f24989a3-4478-4e69-aef5-a8658535b9d4 · outbound

This paper cites AI Risk Management Framework : AI RMF (1.0), January 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Risk Management Framework : AI RMF (1.0), January 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.817088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.234922Z digest=sha256:1c935501ce0c0aa7a2ef4aa47e4c17a7b87524c1fb2a67f3e4353c06d4c48728

Observation 94250e00-55ce-41fe-82d0-0e93f89049fe · outbound

This paper cites Openai system card: December 2024, 2024.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Openai system card: December 2024, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.799205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.239688Z digest=sha256:a3669445e7701022503e8180c8962fa7c50dc2e24c8296a59ab49d50f41c247f

Observation 11943eb2-7684-42a4-9402-04f4fe61a795 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.244674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.244674Z digest=sha256:df506a098e8be11a03ac92926efe7dc500089204479b01c5b7992b1f1e3ffac2

Observation 45cc7030-3d76-4a04-a0b8-f5b5f5e44703 · outbound

This paper cites Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.249627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.249627Z digest=sha256:9fbb2f7f3cd3da5d0e4c0dcbc4f502670286ac43b64910b3ba03398b13e40bf4

Observation a9921514-bedd-4a3d-a40f-c3ecc1d08f61 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.254462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.254462Z digest=sha256:cfa248b357055c004fcc1b803af73f2d92dba959f623fabba94ba0fe23e6afa0

Observation 7657da45-83b4-4452-96cc-aaf6b36efad7 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep, 2024 a.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Safety alignment should be made more than just a few tokens deep, 2024 a

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.783616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.259603Z digest=sha256:03ec745f7101f8f2d26984439128d36ccb18cda3e71a645ecb9e70150e35fda6

Observation 46082899-8732-4da9-ad97-7c6104d9b44c · outbound

This paper cites On Evaluating the Durability of Safeguards for Open-Weight LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On Evaluating the Durability of Safeguards for Open-Weight LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.264207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.264207Z digest=sha256:d87172161864ed670862929f4a63e6c6824c00e0e7127df86796051e08fc0ae8

Observation 46e855c5-3c66-40d2-b412-a59fd126938b · outbound

This paper cites D., Xu, P., Honigsberg, C., and Ho, D.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities D., Xu, P., Honigsberg, C., and Ho, D

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.767226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.269083Z digest=sha256:aa23491030edf759f075d473a26497a2aec49cd7a0ca735fa35ad65a99f206d4

Observation 02e2e5a7-928d-45c4-b7d0-34193dd1bf67 · outbound

This paper cites Open Problems in Technical AI Governance.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Technical AI Governance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.273714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.273714Z digest=sha256:a169ec8e0be4c1ef53505655e6cec455f2205a56a28feac7d7319f55cabd4bc5

Observation d1c44e37-5f68-4494-8618-560ad26d944d · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.278636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.278636Z digest=sha256:358c00e7ab3e9f5e9fcb47c13b2a242bac0552fb4e9b9297b8e9e7a467912a38

Observation dc0996cd-9a84-41e8-95b7-8e63d29f3799 · outbound

This paper cites Fast Adversarial Attacks on Language Models In One GPU Minute.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fast Adversarial Attacks on Language Models In One GPU Minute

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.283768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.283768Z digest=sha256:9c6a9461d3f10d158421a99f768ad3361be2d1548bc7817642d43342ff03378e

Observation 494f9dc6-20ee-48f4-825c-e56e317b936b · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.749567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.288872Z digest=sha256:d451024373ffc6d8f9360289c66d6697e76b05704b2b9717403d011a7a96a4db

Observation 77db4e86-0d51-49ef-b89a-be27da8b1815 · outbound

This paper cites Towards best practices in AGI safety and governance: A survey of expert opinion.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards best practices in AGI safety and governance: A survey of expert opinion

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.293676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.293676Z digest=sha256:61693524e893c4f47de370cc0f92a79a920fabb7d9aef039c8229dab9c93a6ef

Observation 6ccc3f80-0418-4903-a57f-661b4c4fc24c · outbound

This paper cites Adversarial attacks and defenses in large language models: Old and new threats.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial attacks and defenses in large language models: Old and new threats

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.730884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.299109Z digest=sha256:3ae12700e95b510b4f6d5b4e7a94889c7ff308263782eeb9bc570ca5c74e82e2

Observation ad40a32b-3f19-4bfc-b54c-e45792a8415a · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.304017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.304017Z digest=sha256:2488f61268544f8e7536214e09d02fbd6cfe30a0befd26f404633867b65e42e3

Observation af0cb92b-f7bd-42ad-9c34-e3a1ece7f36f · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.309285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.309285Z digest=sha256:44b0be47e9e8f97121d380946a8009baaac0fa5f35625403d703aa226e5dc69a

Observation 5b9b18c3-6020-455f-9cdc-8ab91d420ad9 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.314358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.314358Z digest=sha256:c07843be544235f414286d3395fc97bd3eb95f6e221fc2ba044566de134146cb

Observation e882aafa-ceab-4c62-9b51-c96884513ba0 · outbound

This paper cites Model evaluation for extreme risks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Model evaluation for extreme risks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.319250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.319250Z digest=sha256:ffffe82f439391cdd20eb099ad555a6e331f19db21495e891f8f3cc86e26640b

Observation 04b13167-b951-403d-9833-4c2ac64427dc · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.324252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.324252Z digest=sha256:08822569b91b25fdaa9b0c130d8f8da06684b69c38e46190a447b5d417e331ff

Observation 8b1f053a-be81-4969-aa6a-d73c134f95ac · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.329073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.329073Z digest=sha256:8783f50514936aafb9c13062f4c51336f6b0cfa94c16c432a0714e760117e11a

Observation f6df5fa0-756f-4281-87e5-1492cc8cc155 · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.714557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.334078Z digest=sha256:f3cedede2042af168868935049c43e58dd09130dd4b1b43bf6ecb133c2a5a070

Observation 3b01b132-ce61-46a2-aedf-4dcea52fdc81 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A StrongREJECT for Empty Jailbreaks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.338954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.338954Z digest=sha256:77ade86b30923ecd1c06eea6bea039808fd80e61772fce49828f0467684ed190

Observation 1d3ddeba-b616-47bb-bed5-613a2bf254ed · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Simple and Effective Pruning Approach for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.343866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.343866Z digest=sha256:8017c9c90b072f9dae1cadc8a33671df8bb42a6c708edc605355b6ce6beab845

Observation 9a2486f0-3e1b-4f9f-b395-d26a1835bd4d · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.348837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.348837Z digest=sha256:882107e23903da4986fae0addf477ab2d22df6fba7f0a1d2b282382238d26a60

Observation aa727472-dfef-42ec-84b1-83bb5e94f6d2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.353590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.353590Z digest=sha256:7763f735d7db3739329cd4c5d008fd9030e92e6d462d5f3522f0a06d2d35c47d

Observation 1dfc3555-196d-425e-833a-f95e4ee86445 · outbound

This paper cites A pro-innovation approach to AI regulation.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A pro-innovation approach to AI regulation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.698068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.358517Z digest=sha256:f8a8984d2812a0110addf043381a8b3a5057ed8313830d5bc55f0ee2675c6822

Observation 1876fddd-203d-463e-839f-4af8685db792 · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.363084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.363084Z digest=sha256:637c2c38e73737bfd14aaa51fb5c6e60e8c934d2c7a93779debc909f1de3cf1e

Observation dd41a068-d301-48c4-a177-4c4270a48e03 · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.367750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.367750Z digest=sha256:16b9748ceba8b5d782bafe27b66b9c264b88b6403c8b64a9346f7327123122a8

Observation 0bf8c76c-ecb1-4d01-a8c8-a7a07a00873c · outbound

This paper cites Efficient Adversarial Training in LLMs with Continuous Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Efficient Adversarial Training in LLMs with Continuous Attacks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.372451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.372451Z digest=sha256:86d372f1c5872948a31c74e9ea653ea33a009ad167d684c6828ba2f9c1526201

Observation da6e0ecc-79bb-47cc-80ab-863b8edc129c · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.377658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.377658Z digest=sha256:38d94b9c1bceaa1bf27f060c16064127158f34fc0bc3d541921006960412274d

Observation fc9df143-d719-4af9-8b61-8e95f72f1ce9 · outbound

This paper cites On the vulnerability of safety alignment in open-access llms.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On the vulnerability of safety alignment in open-access llms

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.680996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.382845Z digest=sha256:05c324a37f2657d1e6783331c2a991b1f0f9d101988c09a15a34ea383347f92b

Observation c0b52edb-205f-49c1-879f-be6937fc0b5e · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.387716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.387716Z digest=sha256:6e94e247823a1d1f07ba132b0ad799d4d123a98feb4b0ab4bcc095b36db1e19a

Observation bc50e58e-37ee-492d-bc8c-52347c806179 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Low-Resource Languages Jailbreak GPT-4

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.392539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.392539Z digest=sha256:2f9206cb32be2466f2936c112f0b0557b1d096bff6a12127222e5b8e1101694c

Observation aefc68dd-f5a7-4218-8fe3-77f7c6f34395 · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.397309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.397309Z digest=sha256:3b51b4b9e32fd022caa68ef5a9cd0fc2a1e7e9480a86d3554fe29747fa763bca

Observation ad4eb8ad-9be3-49ed-bbd1-b106c3b322d5 · outbound

This paper cites BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.402376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.402376Z digest=sha256:774ed3e45f1a14589948005ede5bd66ce30a9ccc4d8fb045fe3caa17abb35ee0

Observation 75cfd0a4-16c1-4066-8ef6-5fb86992c319 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.407456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.407456Z digest=sha256:cd4b4014219c15e79ba509c26e0c472dc99585f905290a5509f5511a9c021e79

Observation f3323a75-f9ad-4ce3-b497-b1bad5992b69 · outbound

This paper cites Adversarial machine learning in latent representations of neural networks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial machine learning in latent representations of neural networks

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-08-09T14:47:24.387209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.412546Z digest=sha256:b831e777a07c4d4e53ba6855ff26d1fa4126a4ea1c22e61e24b7cad4c1dc68d8

Observation 73905ea9-a758-443f-8657-588ec377ee6c · outbound

This paper cites Catastrophic Failure of LLM Unlearning via Quantization.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Catastrophic Failure of LLM Unlearning via Quantization

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.422836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.422836Z digest=sha256:2d14f3875cbb25ed281168d2e8d3d4025f2d303bc3f99a89e949c86420df3dd9

Observation c52482b9-2f28-4048-80f8-85433775e0de · outbound

This paper cites A Survey of Recent Backdoor Attacks and Defenses in Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Survey of Recent Backdoor Attacks and Defenses in Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.427626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.427626Z digest=sha256:ccb7d4496a1b2cb7a93c8323a4c9ad9e6a3aafb94669fec9205b7fb9db0ae439

Observation 0f86f630-1d5a-440f-a596-014d8a74a2cc · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.432963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.432963Z digest=sha256:d53f466ff09781a93fe9fa6b24e9241304521ddffed0f02feb051a998970d2de

Observation d94a0715-f578-490c-8a47-ec5e68464ee1 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Engineering: A Top-Down Approach to AI Transparency

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.438083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.438083Z digest=sha256:24f2f2f58da38626aee4c3bd3a6396344dce13a81f7dd0c0c865a4f9374fcd0c

Observation c0a5011b-8beb-48f3-b3fd-c6ba8f62a5b6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.442928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.442928Z digest=sha256:f7c2f415ab27706cfb75ae58c7cd8028f77e98e8bd62aab197d25223fecf8fa0

Observation 73f8f072-4083-4492-8bce-0cafd19c6284 · outbound

This paper cites Improving alignment and robustness with circuit breakers.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Improving alignment and robustness with circuit breakers

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.664850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.448064Z digest=sha256:356e6e45ca024e0cf12533c72897b7beae02fcd6908f9c63ff382e2ec551589e

Pith citing papers

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · inbound

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods cites this paper.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:cc695d8c6f4f4ed0289dceca354ae0d12fb6dc4bb4971ebb98ee9b086bd9d9f8

Observation 5b691c14-d0ac-4408-84b3-1a983d3c9287 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.212266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.212266Z digest=sha256:ad66be2da3fb0394a644646d61c4bc3d8b7f6750b400e35692704e94dc608a26

Observation 0abc5024-ef5d-4655-95da-9c9b68ade291 · inbound

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning cites this paper.

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:46:17.058220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T10:44:53.516653Z digest=sha256:828d7b955191f53050af7e660a6f08417f8ef67eca94210d2232d3310b58bdd3

Observation 9195ab00-e824-41a4-9ef9-dd9eb8448c61 · inbound

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories cites this paper.

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:41:33.209039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:41:33.209039Z digest=sha256:3f24e7d78755c610dc91757a620a475a564794a2dd7f9d485b198917a4cd7a4e

Observation 5ba3b8fb-a6c3-43a7-8dbe-1d7aa8abb55a · inbound

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns cites this paper.

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:59:01.358929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:59:01.358929Z digest=sha256:a61e719205681241e7c35792c16e855f1b801894f1f81a466a5e474c048d9873

Observation 735466bb-16e9-47f0-ad25-d74394d5c890 · inbound

Operationalising the Superficial Alignment Hypothesis via Task Complexity cites this paper.

Operationalising the Superficial Alignment Hypothesis via Task Complexity Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:49:13.999922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:49:13.999922Z digest=sha256:66445fc81131fc283451d7d77c18795772b61e14a526195e0a22ab8022af8c74

Observation 9f63ce60-b49b-432b-94be-3091cd96bea2 · inbound

An Independent Safety Evaluation of Kimi K2.5 cites this paper.

An Independent Safety Evaluation of Kimi K2.5 Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:43:11.577719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:38:18.674355Z digest=sha256:ff24a41d7fb9c316a1a6cd003ef1b70c46af832bb75eeef16489ed8f4dc0072c

Observation ec539424-962b-4ec3-b385-e3a4912b5b1d · inbound

Is your algorithm unlearning or untraining? cites this paper.

Is your algorithm unlearning or untraining? Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.104808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:06:08.962042Z digest=sha256:4385bb72e0eddc9821e66675642c6196dc3e8f5bff8d4610be22a8ed7ccb1482