Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T07:33:43.015966Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.02714.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T07:33:43.015966Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 887ceaa0-8dd3-496b-932b-64e68e3d5505 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale An embarrassingly simple defense against LLM abliteration attacks.arXiv preprint arXiv:2505.19056,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2caab62-2233-4b1d-a14f-600eca1074a7 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24a3ae3-6c27-42ba-9602-6c61124f12f3 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7307f2ab-2956-4622-8719-1bc95332bb44 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A granular study of safety pretraining under model abliteration.arXiv preprint arXiv:2510.02768,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d563e81-3771-4fdf-aec8-a76896277c0a · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Measuring Massive Multitask Language Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82e3692-efa1-40cc-848e-a6c109dfb1c5 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Kimi K2: Open Agentic Intelligence
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739358e2-e78f-449e-9bf9-e17a8ffacc08 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150f3f6e-8b64-4988-9ed7-8e975ea53f56 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe61b63-5eb0-4afc-9559-ec0e545b0e51 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03690b0c-a1d1-4f51-bcea-7e35b029ff2b · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95490576-296c-4b4b-a86d-91eaf42335f7 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721cc18a-5a00-4481-8e61-89bf69f1ccc2 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Steering Llama 2 via Contrastive Activation Addition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9f8f0d-291f-4eac-82f6-f3395d34cc87 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 431f37fd-a722-42b1-958a-3c8766bfe893 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Linear Representations of Sentiment in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c46c56-fc1f-4f6b-9304-b67067814ca3 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Steering Language Models With Activation Engineering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aeee549-91e6-40e4-a8b4-d131e713f3eb · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23f879c-557c-4abf-a2e4-8c4812be1ddb · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277b9721-b613-46f1-822d-794addd4ad76 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Comparative analysis of LLM abliteration methods: A cross-architecture evaluation.arXiv preprint arXiv:2512.13655,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 356a6d6b-3a3b-4480-ba9c-7a016727c5f6 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Representation Engineering: A Top-Down Approach to AI Transparency
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f246d2d2-49a6-4432-b6ad-5b21e3225947 · outbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.