Pith. sign in

Paper Citation Record · LEDGER

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2301.04709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.04709 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:44.215403Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

10
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f2e48289-ffce-497a-9a7b-aa5db8fe4d98 · inbound

Localizing Model Behavior with Path Patching cites this paper.

Localizing Model Behavior with Path Patching Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:37.803414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T19:38:37.751487Z digest=sha256:779b13bfb74095cf41a5ef8b305dc92aa30aa46f85fd2999d3fe11bc059eb866

Observation ac85d2ad-2034-4509-b7ff-a108d493aa27 · inbound

Linear Representations of Sentiment in Large Language Models cites this paper.

Linear Representations of Sentiment in Large Language Models Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:43:06.462170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T12:43:06.361170Z digest=sha256:efb0baa70efa823b64962fd2bccee821420eaccc40a3d3df1e478c63826314f7

Observation 072600d1-7c4a-4cc7-98f8-874fb6884ae6 · inbound

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models cites this paper.

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:15:10.703650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T13:15:10.632115Z digest=sha256:c204c38be50fa29aed83b79375cb63376b92136557d35c9f46b9269de8bbe106

Observation e44f166a-e856-466f-9106-7dacbddeacb6 · inbound

Factored space models: Towards causality between levels of abstraction cites this paper.

Factored space models: Towards causality between levels of abstraction Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:28:45.301776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:28:45.301776Z digest=sha256:83343b4916a8c28e8147116508cd3deee8a1d955b4a63a89dd7c6b7f03368311

Observation 07fe9ccb-2b20-411f-ba91-3906b2393cb6 · inbound

Removing Spurious Correlation from Neural Network Interpretations cites this paper.

Removing Spurious Correlation from Neural Network Interpretations Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:03:10.792600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:03:10.792600Z digest=sha256:ed4e96fb1b543e8449ce7e6048508a543fd27bd8c28a7db4ca387cd823cd43c8

Observation a94c6dd8-759c-4b98-b19c-1e391c6f6b8e · inbound

What is causal about causal models and representations? cites this paper.

What is causal about causal models and representations? Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T20:38:36.470902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:38:36.470902Z digest=sha256:1713b64427a94b268366571352be3d559fa6bfe42bb4cbe5da730458fe4b27b7

Observation 0d7cd494-b1de-4b9b-b4bd-199414ac0a7c · inbound

MIB: A Mechanistic Interpretability Benchmark cites this paper.

MIB: A Mechanistic Interpretability Benchmark Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.215403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.215403Z digest=sha256:03915471206e6f433039e1820d35b71a9c2916fa9193e3e00a000641ae1b20c0

Observation 22aea982-0987-44f7-ab5b-592ababe55dd · inbound

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii cites this paper.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.733991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.733991Z digest=sha256:bb257df63e06b21a497365b2dabbaa912ee1e125cd9c9b3049d02d080c70108c

Observation 831539e5-aef9-43d1-9f5a-1fa53e94a015 · inbound

$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks cites this paper.

$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:57.678412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:57.678412Z digest=sha256:a8842bd5cc244f8bd0d3fe405c2ae54a7d272aa7936842d6305ded9e3e83d0ac

Observation 60e2b358-386b-40d1-b04e-197bf857f57f · inbound

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors cites this paper.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.583205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.583205Z digest=sha256:76190c294922d8896b191ab333602e28fb8ea8438cfd1ae241676855199d58db

Observation 8240af5a-2aed-4c27-985e-c4e89b35704f · inbound

Explaining Neural Networks with Reasons cites this paper.

Explaining Neural Networks with Reasons Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:40.162144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:40.162144Z digest=sha256:7ce151e9ce387df45c1fca10a97fec47c74da20d5aeabd0dbde6eeec0ac61133

Observation ed653d77-400e-476f-b68f-b3871cd5a005 · inbound

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs cites this paper.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.147887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.147887Z digest=sha256:8029dd1e4546701951570a2527f1159eb8fd19290cb686b0d672cc066066dc12

Observation 60b65955-83ca-44f0-98ca-bf7a822f87cd · inbound

How Do Transformers Learn Variable Binding in Symbolic Programs? cites this paper.

How Do Transformers Learn Variable Binding in Symbolic Programs? Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:07.811798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:07.811798Z digest=sha256:17f062cff7c41adbbee7edcf2d124e7d66e504207170a402aaa9b8738ad7e5fd

Observation a99928f2-fe36-4677-b909-72e3eae049fc · inbound

Identifying a Circuit for Verb Conjugation in GPT-2 cites this paper.

Identifying a Circuit for Verb Conjugation in GPT-2 Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:46.623506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:16:46.623506Z digest=sha256:3419735cba0db323323ca27c0da46275f0f574946ae5f57d85182f4e83b654b0

Observation 88fa77e4-1605-455c-a09d-0b79149d539c · inbound

Identifiability in Causal Abstractions: A Hierarchy of Criteria cites this paper.

Identifiability in Causal Abstractions: A Hierarchy of Criteria Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:22:28.725891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:22:28.725891Z digest=sha256:be0a5adeb3f3c38eab4b094d54a9a2fc6a17a6a99d85c9c145ff767726c46129

Observation 0e43d9ef-db7d-46cd-aa37-b05d7851210e · inbound

Can Interpretation Predict Behavior on Unseen Data? cites this paper.

Can Interpretation Predict Behavior on Unseen Data? Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:28.638962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:28.638962Z digest=sha256:f301e07321b0874432c64ba7272556bb856df6fb1def9ac757c9b66f1616a1d1

Observation 4ecdeea4-ca1e-4230-8825-1855bc0ca1db · inbound

Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies cites this paper.

Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:37.348803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:37.348803Z digest=sha256:ef571df99f6b21b10e30d195e50dcfee1cabd595335705879f4c190343a5d4a1

Observation d7dee443-348d-4baa-98ad-9c9fe9f2a572 · inbound

How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding cites this paper.

How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:12.926237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:21:12.926237Z digest=sha256:1a018f5451a9d7f13305efa79ede9526cbe2cf5adcfe2d9c6ba966aa1155dd69

Observation 03fbbe87-c027-441e-a4b8-17876f48c024 · inbound

LLMs Should Not Yet Be Credited with Decision Explanation cites this paper.

LLMs Should Not Yet Be Credited with Decision Explanation Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:06:33.196776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T18:41:24.552939Z digest=sha256:5232589cd1580699736ff55f0f28765ebe756ba21908807b174dd53d407a8fa6

Observation 38219848-7677-43dd-aed7-9a0cbec2407f · inbound

Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction cites this paper.

Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:36.290694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T19:12:30.627633Z digest=sha256:74196b049eed7b0eab5df5a010cb44e7edf792e9f600f4b8dec3755bd7cef203

Observation 9395ec40-e26a-4326-91d3-68fa7c649b7c · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.264696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:e2ac23a4c304be33b7b5e5a6f6d132e020b889249a81addb2cc15ded49ceb28d

Observation dcc8bb27-776b-4289-895d-4364afb81e32 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.151029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:36ccc082faa4ab92acd643040ebc8d3a3ecd762217ed46c272b188429e4fc505

Observation 540a2dd8-b6c1-4a46-b7d8-b72425f6549b · inbound

From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks cites this paper.

From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:07:41.081855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T16:06:05.508610Z digest=sha256:bb5fc03f39f65107bf088acca9aedc4310e880eb8d585b2244e4424c68ae1e7a

Observation e0173868-2849-46b3-9fed-ba9a2a8aec7b · inbound

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express cites this paper.

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:13:59.318138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:13:32.378862Z digest=sha256:3a435c6ad7992915804e16affe9e7c6ab3b472f32b128f3a0e67954221927986

Observation 82492bf6-f21a-4999-941f-563b442d697d · inbound

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English cites this paper.

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.895670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T18:44:23.296625Z digest=sha256:0265a8d81b3b00b59f172292503b1221d9023a17c206033aaa78f5f45d67d7d2

Observation 9dbb2565-b89d-4eed-9a79-7e1331e76973 · inbound

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English cites this paper.

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:14:20.387389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:14:20.387389Z digest=sha256:82a25b15f1e7065371b2a0e8e60659f32d9a446d8b0080a6a07871c2c4891e36

Observation c934caf3-ffe0-40ac-961f-61d6a5a5394a · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.137704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:634b27fdfe8383dbc4f40a282f7b4f43aca57f84ed9d788d668129afb6fa9517

Observation 7b896aa9-f0cd-4ff4-8612-394cf17fddc8 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:87c3c32b37930ce11862fd6d8011d75acf87dac1a8c275548a147deb3b243b33

Observation f2455939-37e6-42cb-8548-afb76926aa3b · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.287226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:7c67e1ec5d2f875965a41c78d3c56aeed0455005e3272933859fea631041c554

Observation 6204ff38-c9a2-4847-8cf6-6787df2f3870 · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.074500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:86d6aacc76318d4014465d94800bb96eee8639214c26b4a513c1f52153bb411d

Observation 201b7313-9d69-4b6e-b126-58c5af926390 · inbound

XtrAIn: Training-Guided Occlusion for Feature Attribution cites this paper.

XtrAIn: Training-Guided Occlusion for Feature Attribution Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.382726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T13:37:35.691503Z digest=sha256:1c33bb40c75112934cc7b95948e822ebf81aee09aa0404bc63ffe5162a58b0aa

Observation 832bb100-a667-4f0d-adbb-bf5ed7e9d6c0 · inbound

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks cites this paper.

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.217596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:49:42.069456Z digest=sha256:b5a43b146090b8757bccd0b4aa5bf370777c1f1d4dbe525987c0b3f8ccbe1a6c

Observation c6fe8413-40de-4ea3-aecc-856b9dbde6f4 · inbound

Steering Vision-Language Models with Joint Sparse Autoencoders cites this paper.

Steering Vision-Language Models with Joint Sparse Autoencoders Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:07.779400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-25T20:56:57.716246Z digest=sha256:7f994ea97072ed8b5133fada17271b3f820d660d3d80c667e0bff218bf9ce282

Observation 51b25163-ad74-4382-87a1-03216aaa2ef6 · inbound

Safety from Honesty in a Disinterested AI Predictor cites this paper.

Safety from Honesty in a Disinterested AI Predictor Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.196264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T06:53:19.737408Z digest=sha256:29720b9ee5e294f433b576b813b315545f0f685518d32092aeaef1d309aad38e

Observation 890e0878-d980-4dd5-9b5e-32e7b8e7c48c · inbound

Safety from Honesty in a Disinterested AI Predictor cites this paper.

Safety from Honesty in a Disinterested AI Predictor Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T17:02:19.189814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T17:02:19.189814Z digest=sha256:7afaba5f4fca11c6ff268bf4c0960f6b14f74b95f72c7163ce16065da0dc66e6

Observation 891db3d4-d0ba-45a8-b03d-7375297297fd · inbound

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects cites this paper.

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:03:55.219054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:03:55.219054Z digest=sha256:daa3ab1434c6b3eef2a1150badff527b517f54bb7845f237aae137a361bab72a