Pith. sign in

Paper Citation Record · LEDGER

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2507.09406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09406 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:51.737979Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 12d5f8d8-3aa0-4fc5-80ad-1915e43d841c · outbound

This paper cites Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.639435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.639435Z digest=sha256:be99919cebb9d7e53c9bd1ca50bf691bded9d49fabaf7dc10d29e0f777b31a04

Observation 7ac2f7d0-f8ea-46b0-9f18-77c51008f7e5 · outbound

This paper cites Advances in Tabulating Carmichael Numbers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Advances in Tabulating Carmichael Numbers

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:52.053060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:00:51.660510Z digest=sha256:947a0d43af75e2e815fd3f445a0090739c8b13908df645c13ab4629cdfefed9a

Observation 4e4a4539-e5c8-4b12-b1a2-387df5881be0 · outbound

This paper cites Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.667312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.667312Z digest=sha256:3fd06a453f926c2094ee13b7ef708f566d389786fa1a0f6bda86a8cb2c47f70c

Observation dd598a0e-193b-40c0-8e9c-e236ded13ec8 · outbound

This paper cites Generative Adversarial Transformers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Generative Adversarial Transformers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:00:52.006017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:00:51.674334Z digest=sha256:de240493f2387fe256b2d25d046bd8b38fbb7a546f1121b3f0447193298a1bfd

Observation fccdbb04-1984-4470-9093-8fde84001115 · outbound

This paper cites Wav-KAN: Wavelet Kolmogorov-Arnold Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Wav-KAN: Wavelet Kolmogorov-Arnold Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.691306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.691306Z digest=sha256:f12e76db6730781071aa311f038786263a67c44f84addece164e7c9f6fd8cdd2

Observation b03f4675-9e8a-470a-af2a-b3a9d1a67ef9 · outbound

This paper cites Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.937745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:00:51.699652Z digest=sha256:5a890ddc0e9b33ca1cfd992bf7811e0f6e26a13f8c0123ae746768d79373e141

Observation d94b862f-9d0a-46a8-8e48-25a4526a7381 · outbound

This paper cites When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.707342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.707342Z digest=sha256:a64ac28344c92509d6648fcbe058e6f1479b4051f570895a0b3454f04a9d4a71

Observation 76678cc1-2453-42e8-9002-4947eb8d2b69 · outbound

This paper cites Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.868458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:00:51.720004Z digest=sha256:73f1a53031433e325e1d8b86baaa596932e6fde1a60eedb03a6be98ec9bd6bae

Observation 44b6f56d-e98c-43a7-8ad2-cc33bdc55e65 · outbound

This paper cites Guidelines to Develop Trustworthy Conversational Agents for Children.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Guidelines to Develop Trustworthy Conversational Agents for Children

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.798306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:00:51.737979Z digest=sha256:1f0f6638f18a8289b64dccd1e99885ac1cf934b9f250af68a978b143f0124838

Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.725380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.725380Z digest=sha256:999ffb8e80f286ba9ccae05110931e8094d327fbe236e410586865bfc9393883

Observation 77779af3-9269-45de-8d9f-813008f96dae · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.683998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.683998Z digest=sha256:d32d2063eaf622b46862b139081cbcfd5bb1244d46b463486988f08ec15c3a22

Observation db60480d-bbfd-433e-994d-e15860c42690 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.731106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.731106Z digest=sha256:b8993f9f9336bd494b9fc9e47f1ad8be7a18cae9c0b324ee867538ecfd6f7740

Observation 524799c9-cfad-4af9-8e34-35d4c6851b49 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Mechanistic Interpretability for AI Safety -- A Review

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.646078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.646078Z digest=sha256:f92be5393dc22c9c133c172647ece2f8c3d46c62ee28842c4a3355375c853e9e

Observation cff18f5d-3676-41d1-9e20-d22d8d0d5311 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.654377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.654377Z digest=sha256:f636571c1cfafc11d2b945749125808d380876e1b547823d0e781ee7f8359ead

Observation b23b4ea6-c812-4b64-8cbd-3fec2f95ea8b · outbound

This paper cites Attribution Patching Outperforms Automated Circuit Discovery.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Attribution Patching Outperforms Automated Circuit Discovery

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.713937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.713937Z digest=sha256:a3d188f9ab3775872a13e5da17070a9808431cc8d091de8b05d8852d94b99333

Pith citing papers

Observation 98504e4f-8bf1-491c-9c42-e344c02384b3 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.785630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:f67fc650783572af9602692b557d28fcdf066fe54dc78c359b0a2c7878236f28

Observation dd2e7047-a80d-4aca-bb02-c463207269ec · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.976229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:e1adac4d1505b8aa1888cda6a9d8bb5cc95d01009f718ac5caa7bbd2963d34ad