Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.448064Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 8 inbound Pith citation observations for arXiv:2502.05209.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.448064Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:06.683217Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T10:46:17.055047Z
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf1f4201-9b52-47eb-a0b6-d553ec611530 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183e7ffa-d049-4cfb-b8cf-0c3e179c6892 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b1af65-f276-4185-8a02-8c85bf245625 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f4dc67-561b-4b2f-a9fa-dbd1744c6fb6 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Many-shot jailbreaking
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e72b8b-62a8-481d-8271-b19d296bebaa · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning in large language models via activation projections, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60765c1-c7ba-4511-b905-6f60d5639e29 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refusal in Language Models Is Mediated by a Single Direction
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633bc806-8686-4c0c-bb0e-16e5805b4fdd · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d4dc73-3dd6-4f1b-be7c-ac44238e7d1c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Machine Unlearning for AI Safety
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fcfcd8-651e-4314-9848-00fa399ccd0d · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f9795a-a64c-442d-95fd-3ba18fbca45d · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a305ddd1-337a-4a51-a1d5-5426e98c3f80 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2320bc-7954-45f0-8146-5a6b68311de4 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A., Jagielski, M., Gao, I., Koh, P
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375996cf-fb2f-4c96-8acc-b63e4b626094 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b042a6f-9ed9-45de-ada7-14c98f4f361c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b04b81-e9bb-4b32-bb5f-815e9bd6bf24 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Interim Measures for the Management of Generative Artificial Intelligence Services , 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 037a7571-e26d-4a03-8e60-13c6c5568b60 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities G., Islam, M
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd04bde-0107-41eb-b320-f5373135624a · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities O., and Nilsson, F
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f515d96-f150-4a66-87a0-e43671afe57a · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Do Unlearning Methods Remove Information from Language Model Weights?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e625dda2-cebb-471c-a6ef-c43baa3c9637 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafdf18d-29b6-43e4-b8a6-8cf8436227fa · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The eu artificial intelligence act
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feea78e1-b98f-41dc-94a2-5ebc51bc4061 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Scaling Laws for Adversarial Attacks on Language Model Activations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9d2a9f-2592-4aa0-9ac8-804cadcd21c5 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards a science of ai evaluations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1812b8-3cba-45e1-ab28-f5d46d864bf3 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Erasing Conceptual Knowledge from Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef386264-db9e-4fb4-bfb3-572786246e45 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Stress-Testing Capability Elicitation With Password-Locked Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced76700-4f74-427c-98f4-f4fd392bcd53 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Cascade: Exploring hierarchical inference in language models, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122f75cd-0c4e-4a4f-b5d9-096bb7b5eeee · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., Haghtalab, N., and Steinhardt, J
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 859381dd-9d09-4cc8-ab4a-ecce32ab1439 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Measuring Massive Multitask Language Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b4633c-ab77-48bb-be91-52cd0af80adc · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac493279-4922-49e3-a797-b466f436aed1 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA: Low-Rank Adaptation of Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c267204e-d9f5-4a06-a2ef-70836602af39 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd47542-8160-46b8-a070-04a1f37247bd · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bad5977-5ee9-49b1-a799-7ab9661c9872 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853c3ab5-e289-40be-89ea-4a61d4329711 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language models resist alignment, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation beaf5ec8-3b43-4564-b970-5b3fa8a64530 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c6eb7ce-c036-404c-b7b8-55c92326ee63 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Act on the protection of personal information, 2025
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8bb16e20-1d1b-44ef-82f2-e26f100e351b · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43acf8e1-359d-47e2-b21a-ca64edf4fd8d · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85b16c9-7c66-4b34-93ef-ddfa98ce1114 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c86131-280a-44a2-b768-ca80477cde25 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7569bbfc-337a-4e57-827e-c21e17128f0b · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c5d3e0-6632-47ea-a285-1a98c0d9d47c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Continual learning and private unlearning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a67d7ca-f77f-4652-8175-fc201462abfa · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Rethinking Machine Unlearning for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153ff993-0b35-4b9e-a6a0-c6be10f7c39e · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Threats, Attacks, and Defenses in Machine Unlearning: A Survey
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa340c9-2760-493a-9f3e-6fed376a1db0 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Large Language Models Relearn Removed Concepts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a68a4a-4299-4fca-ba6e-6f07c0521583 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities and Rimsky, N
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 52289a63-379a-44f5-b904-0db3406e363f · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities An Adversarial Perspective on Machine Unlearning for AI Safety
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24c5f8c-9705-45b0-b204-b9a06d97a898 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Eight Methods to Evaluate Robust Unlearning in LLMs
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce66b5b-93f5-4eb5-883c-4f0784adca3a · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Pointer sentinel mixture models, 2016
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9b5cb3-561d-4abb-83f5-92739136e69a · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f24989a3-4478-4e69-aef5-a8658535b9d4 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Risk Management Framework : AI RMF (1.0), January 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94250e00-55ce-41fe-82d0-0e93f89049fe · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Openai system card: December 2024, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11943eb2-7684-42a4-9402-04f4fe61a795 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45cc7030-3d76-4a04-a0b8-f5b5f5e44703 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9921514-bedd-4a3d-a40f-c3ecc1d08f61 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7657da45-83b4-4452-96cc-aaf6b36efad7 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Safety alignment should be made more than just a few tokens deep, 2024 a
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46082899-8732-4da9-ad97-7c6104d9b44c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On Evaluating the Durability of Safeguards for Open-Weight LLMs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e855c5-3c66-40d2-b412-a59fd126938b · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities D., Xu, P., Honigsberg, C., and Ho, D
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02e2e5a7-928d-45c4-b7d0-34193dd1bf67 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Technical AI Governance
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c44e37-5f68-4494-8618-560ad26d944d · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Noising: A Defence Mechanism Against Harmful Finetuning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0996cd-9a84-41e8-95b7-8e63d29f3799 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fast Adversarial Attacks on Language Models In One GPU Minute
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494f9dc6-20ee-48f4-825c-e56e317b936b · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77db4e86-0d51-49ef-b89a-be27da8b1815 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards best practices in AGI safety and governance: A survey of expert opinion
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ccc3f80-0418-4903-a57f-661b4c4fc24c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial attacks and defenses in large language models: Old and new threats
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ad40a32b-3f19-4bfc-b54c-e45792a8415a · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af0cb92b-f7bd-42ad-9c34-e3a1ece7f36f · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9b18c3-6020-455f-9cdc-8ab91d420ad9 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e882aafa-ceab-4c62-9b51-c96884513ba0 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Model evaluation for extreme risks
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b13167-b951-403d-9833-4c2ac64427dc · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1f053a-be81-4969-aa6a-d73c134f95ac · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6df5fa0-756f-4281-87e5-1492cc8cc155 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3b01b132-ce61-46a2-aedf-4dcea52fdc81 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A StrongREJECT for Empty Jailbreaks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3ddeba-b616-47bb-bed5-613a2bf254ed · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Simple and Effective Pruning Approach for Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2486f0-3e1b-4f9f-b395-d26a1835bd4d · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Tamper-Resistant Safeguards for Open-Weight LLMs
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa727472-dfef-42ec-84b1-83bb5e94f6d2 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dfc3555-196d-425e-833a-f95e4ee86445 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A pro-innovation approach to AI regulation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1876fddd-203d-463e-839f-4af8685db792 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd41a068-d301-48c4-a177-4c4270a48e03 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf8c76c-ecb1-4d01-a8c8-a7a07a00873c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Efficient Adversarial Training in LLMs with Continuous Attacks
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da6e0ecc-79bb-47cc-80ab-863b8edc129c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9df143-d719-4af9-8b61-8e95f72f1ce9 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On the vulnerability of safety alignment in open-access llms
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0b52edb-205f-49c1-879f-be6937fc0b5e · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc50e58e-37ee-492d-bc8c-52347c806179 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Low-Resource Languages Jailbreak GPT-4
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefc68dd-f5a7-4218-8fe3-77f7c6f34395 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4eb8ad-9be3-49ed-bbd1-b106c3b322d5 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75cfd0a4-16c1-4066-8ef6-5fb86992c319 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Removing RLHF Protections in GPT-4 via Fine-Tuning
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3323a75-f9ad-4ce3-b497-b1bad5992b69 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial machine learning in latent representations of neural networks
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 73905ea9-a758-443f-8657-588ec377ee6c · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Catastrophic Failure of LLM Unlearning via Quantization
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c52482b9-2f28-4048-80f8-85433775e0de · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f86f630-1d5a-440f-a596-014d8a74a2cc · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d94a0715-f578-490c-8a47-ec5e68464ee1 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Engineering: A Top-Down Approach to AI Transparency
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a5011b-8beb-48f3-b3fd-c6ba8f62a5b6 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f8f072-4083-4492-8bce-0cafd19c6284 · outbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Improving alignment and robustness with circuit breakers
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · inbound
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b691c14-d0ac-4408-84b3-1a983d3c9287 · inbound
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0abc5024-ef5d-4655-95da-9c9b68ade291 · inbound
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9195ab00-e824-41a4-9ef9-dd9eb8448c61 · inbound
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba3b8fb-a6c3-43a7-8dbe-1d7aa8abb55a · inbound
Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 735466bb-16e9-47f0-ad25-d74394d5c890 · inbound
Operationalising the Superficial Alignment Hypothesis via Task Complexity Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f63ce60-b49b-432b-94be-3091cd96bea2 · inbound
An Independent Safety Evaluation of Kimi K2.5 Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec539424-962b-4ec3-b385-e3a4912b5b1d · inbound
Is your algorithm unlearning or untraining? Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.