Pith. sign in

Paper Citation Record · LEDGER

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2507.09406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09406 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:51.737979Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 12d5f8d8-3aa0-4fc5-80ad-1915e43d841c · outbound

This paper cites Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.639435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.639435Z digest=sha256:a40f27e3e2edf97ed08b1aeb99c884d05553c1bfd003e45bbeb9a9ca7bf995e2

Observation 7ac2f7d0-f8ea-46b0-9f18-77c51008f7e5 · outbound

This paper cites Advances in Tabulating Carmichael Numbers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Advances in Tabulating Carmichael Numbers

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:52.053060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:00:51.660510Z digest=sha256:b05b47c117282b8308ab6fa28c5b671bb28d5a8c839f4ddb6617a1e68383f496

Observation 4e4a4539-e5c8-4b12-b1a2-387df5881be0 · outbound

This paper cites Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.667312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.667312Z digest=sha256:3fd06a453f926c2094ee13b7ef708f566d389786fa1a0f6bda86a8cb2c47f70c

Observation dd598a0e-193b-40c0-8e9c-e236ded13ec8 · outbound

This paper cites Generative Adversarial Transformers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Generative Adversarial Transformers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:00:52.006017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:00:51.674334Z digest=sha256:7ec0180d36000eb923bf435129412ebc6c109b7999790e50a19c90b90aa0c278

Observation fccdbb04-1984-4470-9093-8fde84001115 · outbound

This paper cites Wav-KAN: Wavelet Kolmogorov-Arnold Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Wav-KAN: Wavelet Kolmogorov-Arnold Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.691306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.691306Z digest=sha256:cb937bc19a3a543bceabac4d6f6b8cc62582c7184a3f5a98dd5682fa9016fb14

Observation b03f4675-9e8a-470a-af2a-b3a9d1a67ef9 · outbound

This paper cites Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.937745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:00:51.699652Z digest=sha256:5ce8127d79d13990302d4bc747e5b2899f8ea0099ebb6af5baf560c03e79a335

Observation d94b862f-9d0a-46a8-8e48-25a4526a7381 · outbound

This paper cites When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.707342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.707342Z digest=sha256:a64ac28344c92509d6648fcbe058e6f1479b4051f570895a0b3454f04a9d4a71

Observation 76678cc1-2453-42e8-9002-4947eb8d2b69 · outbound

This paper cites Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.868458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:00:51.720004Z digest=sha256:988b6a308cede1de3e8cca709bcbbd9166bfc04bcbae90a4fa8c8c5a913d8a32

Observation 44b6f56d-e98c-43a7-8ad2-cc33bdc55e65 · outbound

This paper cites Guidelines to Develop Trustworthy Conversational Agents for Children.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Guidelines to Develop Trustworthy Conversational Agents for Children

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.798306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:00:51.737979Z digest=sha256:0796b23488db98ead6d3460b94e1051a5abab360916078649d486f7d0ae9ac4d

Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.725380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.725380Z digest=sha256:999ffb8e80f286ba9ccae05110931e8094d327fbe236e410586865bfc9393883

Observation 77779af3-9269-45de-8d9f-813008f96dae · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.683998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.683998Z digest=sha256:d32d2063eaf622b46862b139081cbcfd5bb1244d46b463486988f08ec15c3a22

Observation db60480d-bbfd-433e-994d-e15860c42690 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.731106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.731106Z digest=sha256:b8993f9f9336bd494b9fc9e47f1ad8be7a18cae9c0b324ee867538ecfd6f7740

Observation 524799c9-cfad-4af9-8e34-35d4c6851b49 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Mechanistic Interpretability for AI Safety -- A Review

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.646078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.646078Z digest=sha256:f92be5393dc22c9c133c172647ece2f8c3d46c62ee28842c4a3355375c853e9e

Observation cff18f5d-3676-41d1-9e20-d22d8d0d5311 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.654377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.654377Z digest=sha256:f5fe25c38116a61c696bc512a10957b5683ea33bfd5417424a57fed0ab71bd3f

Observation b23b4ea6-c812-4b64-8cbd-3fec2f95ea8b · outbound

This paper cites Attribution Patching Outperforms Automated Circuit Discovery.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Attribution Patching Outperforms Automated Circuit Discovery

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.713937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.713937Z digest=sha256:a3d188f9ab3775872a13e5da17070a9808431cc8d091de8b05d8852d94b99333

Pith citing papers

Observation 98504e4f-8bf1-491c-9c42-e344c02384b3 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.785630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:b4d0875e8f53ba714319db45f81f5190cdb06b51576de6f27414bd9b2d0089e6

Observation dd2e7047-a80d-4aca-bb02-c463207269ec · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.976229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:7b51f839108b133d40eff5abd26ee4acc83dc77f5ce035bca3ebc394d58ba352