Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.05578.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1145ef7f-047a-4ef1-8ea8-a333b10b3283 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85033a13-9694-4ca9-b036-b4c508dc651b · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Dolphin: An uncensored, unbiased language model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a8ccaf23-60e3-4135-81da-5273afaaff09 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Representation Engineering: A Top-Down Approach to AI Transparency
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed92bf25-8632-4f6f-acfa-18a076b6f4ce · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Steering Language Models With Activation Engineering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d430b2db-35fb-49e0-89b2-2b967986b074 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Instructional Fingerprinting of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dba1c4-a689-4636-a7e1-22455d71b97e · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 912bbfdd-f564-45b6-b6f1-af2451286584 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f998b65-7b57-40f4-8e23-58d45a7e8509 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d4c3a31-a753-4cdb-8bf0-9a98a4617f0f · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Sokhansanj
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 36a9fd31-55df-475f-a3f3-8fa8785b160b · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0e3c76-2e48-4f09-8a88-634e2dbaacd8 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Defending Large Language Models Against Attacks With Residual Stream Activation Analysis
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dd234149-1813-4a52-9f04-abdb222f3955 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Safety Layers in Aligned Large Language Models: The Key to LLM Security
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6de393-72fc-4fe3-8aea-7a496d7ddcde · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a315e742-ee52-403d-b15e-bc24df6b0682 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Gonzalez, Hao Zhang, and Ion Stoica
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9fbef769-486b-4882-acad-c3121e68a00b · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706e3b47-05cf-4bee-b850-6f0c68768f80 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis Lo- cating and Editing Factual Associations in GPT
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 87cae2ed-8127-4b6c-a6c9-d1effc757772 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11f04dd4-9a9c-42ba-86b3-792f5256a686 · outbound
Detecting Safety Training Modification in Language Models via Activation Analysis VLLM_ALLOW_INSECURE_SERIALIZATION
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.