Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2401.01967.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:46:29.654391Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 39a753dc-e3e7-4ba2-b3fb-147aff02c503 · inbound
Refusal in Language Models Is Mediated by a Single Direction A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1f7d08c-4030-4d6a-8592-d22be6238419 · inbound
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7646f38f-246d-4cbe-8366-4d53756e1f44 · inbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36f4cbe-ae25-4a6e-bb70-b6b9ea3c6815 · inbound
Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d9fbcc-b2c3-4b1e-9051-d532b70b21de · inbound
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · inbound
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49cceb30-af28-4b3a-b8f0-90f368234dec · inbound
NEAT: Concept driven Neuron Attribution in LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9769c2f-bf46-42c4-b8b9-c147b6fdeca4 · inbound
Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4dd04c29-8bed-4f24-b922-26441193a6e9 · inbound
Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96001828-996e-494c-92ac-e730bc7f142d · inbound
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3569a863-06fd-4737-bdcf-5c53a2b1b5c7 · inbound
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f4cb83e-1551-45f3-b03c-ff7da9755cbe · inbound
Why Do Large Language Models Generate Harmful Content? A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35c5450e-e424-4d72-96eb-05b67664760f · inbound
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9dbf1532-b749-464d-87f9-222186cb1dd1 · inbound
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3dc7a409-db2d-4c2c-bbd3-d78c2cdd19e2 · inbound
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1d6b79df-6825-44b4-9e58-0f800a339e40 · inbound
Tracing Persona Vectors Through LLM Pretraining A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c509664a-987b-47ec-b0eb-42feb21884fc · inbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bcc42c2c-c0f5-4d3d-9cc5-a17265b858cd · inbound
Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac599741-e851-4fce-803a-7f96c139ecba · inbound
Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c6a21310-61a7-4cab-873d-3f68b75b92a5 · inbound
RepSelect: Robust LLM Unlearning via Representation Selectivity A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d45aad08-ea90-46c5-83e3-9014c53f1898 · inbound
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b136081-fd06-4b5c-8f8e-060067cc56a6 · inbound
Tracking Representation Dynamics in Large Language Models with Persistent Homology A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 739358e2-e78f-449e-9bf9-e17a8ffacc08 · inbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365eb461-c285-4511-9f18-379cb93a83d8 · inbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.