Pith. sign in

Paper Citation Record · LEDGER

A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2401.01967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.01967 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:46:29.654391Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39a753dc-e3e7-4ba2-b3fb-147aff02c503 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:47:56.007880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:72dfcdd8a4051b2203986f0f28083c02dfcd8c97fa659fa9252f119a28b18fc9

Observation a1f7d08c-4030-4d6a-8592-d22be6238419 · inbound

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cites this paper.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.654391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.654391Z digest=sha256:69b0af616cd030150a6c6434b75ca62fb659feaaf8b8b73f2e3a7af1c481e3ce

Observation 7646f38f-246d-4cbe-8366-4d53756e1f44 · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.292674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.292674Z digest=sha256:05e090e71d83797a4ae360fa79388ab8e1310cbbd025176a0878fdb5a4202fb2

Observation f36f4cbe-ae25-4a6e-bb70-b6b9ea3c6815 · inbound

Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs cites this paper.

Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:16.715517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:16.715517Z digest=sha256:3a7153b53eb3972acd38de7e0e423b530d661495a8e03b2783aad0301ae79784

Observation c7d9fbcc-b2c3-4b1e-9051-d532b70b21de · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.589127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:063f7a404008a39515940750894d14d2e22c567062cd4ac157c58d8ca48f2c86

Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.735447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.735447Z digest=sha256:95762c08b86355b7ec8dde119bd269403060974218fef9d0b5db56c025d68f8b

Observation 49cceb30-af28-4b3a-b8f0-90f368234dec · inbound

NEAT: Concept driven Neuron Attribution in LLMs cites this paper.

NEAT: Concept driven Neuron Attribution in LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T18:01:19.652668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:01:19.652668Z digest=sha256:26b1f30d0b56907df48a1aa8a96a6dcabc9be466adaccd4b3befd6522e7344bd

Observation d9769c2f-bf46-42c4-b8b9-c147b6fdeca4 · inbound

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight cites this paper.

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:21:23.755068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T13:21:18.699465Z digest=sha256:61ae1913793548a20dee7db5f67ada5536f8e589cb94988179a2aef76dd759b7

Observation 4dd04c29-8bed-4f24-b922-26441193a6e9 · inbound

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models cites this paper.

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:48.897724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:43:48.897724Z digest=sha256:20bff6ede17e198cc40e148267852636c8c7579d7d0b91829c87ecba1e5b0a9d

Observation 96001828-996e-494c-92ac-e730bc7f142d · inbound

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting cites this paper.

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:49:33.106792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:49:33.106792Z digest=sha256:3696e6774cb4f9cd4dbfa72bbb0bccf7eff60033c3c05b1a3f4ccf846b26dd88

Observation 3569a863-06fd-4737-bdcf-5c53a2b1b5c7 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:28.951146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:c9ca82f5ed7aa5f321978edb1e83f2adc40ad8921fc9c6a9a9b6343cb292a436

Observation 9f4cb83e-1551-45f3-b03c-ff7da9755cbe · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:04.435702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:fd4ce063b778daf7a79b299145f6d6acfd9a376bfb83ed1293879582304c6b73

Observation 35c5450e-e424-4d72-96eb-05b67664760f · inbound

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification cites this paper.

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:05:21.995701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:04:13.987667Z digest=sha256:33e6affaaca2bc6a8092c000dfcb45e351c409aa93926e120134e971bc6b2126

Observation 9dbf1532-b749-464d-87f9-222186cb1dd1 · inbound

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs cites this paper.

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.498678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:34:14.310656Z digest=sha256:d256fd3f5c1bfc040c728df2c7b3fe035140ad4c1808de0ec3bbd5f2370ea524

Observation 3dc7a409-db2d-4c2c-bbd3-d78c2cdd19e2 · inbound

Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions cites this paper.

Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:30.228824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:31:40.195348Z digest=sha256:ea2a7ba4ac7b31326ad751dc82dc6992e39ef43ca3a51984af3ffe09044c9db5

Observation 1d6b79df-6825-44b4-9e58-0f800a339e40 · inbound

Tracing Persona Vectors Through LLM Pretraining cites this paper.

Tracing Persona Vectors Through LLM Pretraining A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:29:27.871025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:28:17.086117Z digest=sha256:d9672a7203f60acc324cca0630a806d8a7b37737c2a3be215c0ed94b570d252e

Observation c509664a-987b-47ec-b0eb-42feb21884fc · inbound

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models cites this paper.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.711820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:8e441048f4e7f617384d67e88a7b461a04fe56733fca4d5977ec2cc633fe6593

Observation bcc42c2c-c0f5-4d3d-9cc5-a17265b858cd · inbound

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability cites this paper.

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:11:29.059463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:07:18.198225Z digest=sha256:aeb62d2716bdae3b3aed8420d7892e9aff8617fc8258c1676dd9f18f4e789150

Observation ac599741-e851-4fce-803a-7f96c139ecba · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.949014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:25ed3dbb8144b844dc8f4aaac519d09f5c80e9a1d71d9d08f885b712bafeb957

Observation c6a21310-61a7-4cab-873d-3f68b75b92a5 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.166430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:c999d7fd9f4f39f1c6067ba22e67528365c4d16cee0e4af9320b925c25c70377

Observation d45aad08-ea90-46c5-83e3-9014c53f1898 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.279218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:53c7c2e3c4b39d476a92dfd4edf6b4970cd7dbd4da4c7e108df643997eb2a32b

Observation 4b136081-fd06-4b5c-8f8e-060067cc56a6 · inbound

Tracking Representation Dynamics in Large Language Models with Persistent Homology cites this paper.

Tracking Representation Dynamics in Large Language Models with Persistent Homology A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.225595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:03:07.253265Z digest=sha256:3ac618deff001cb905e26010de8da602969629608c292433e0368beca640dff6

Observation 739358e2-e78f-449e-9bf9-e17a8ffacc08 · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:4c699304a902eade3fc47ba6d8740b1a224fd9fbede6ae9c822bd617733298c1

Observation 365eb461-c285-4511-9f18-379cb93a83d8 · inbound

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates cites this paper.

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:32:04.215688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:32:04.215688Z digest=sha256:e6561f3450efcdfae756191e42967846fa6088d9fb5edc538b17239df5a6defb