Pith. sign in

Paper Citation Record · LEDGER

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

As of 17 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 1 inbound Pith citation observation for arXiv:2505.01372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01372 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:25:06.108473Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:42:26.173782Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact5
  • verified fuzzy33
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7be2a119-13da-4a7a-a1d7-649e3984d9d1 · outbound

This paper cites The Computational Complexity of Circuit Discovery for Inner Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Computational Complexity of Circuit Discovery for Inner Interpretability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.518727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.518727Z digest=sha256:01f8e28fc082a568a616d6d50f68dc2ff3702116f4c523601317a7d8cad319dd

Observation e96be6eb-3233-415f-bede-dba250dc41a7 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.526366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.526366Z digest=sha256:3a4a4f394251ce1cce0212426cb49af0c56296ef1f03888c1d2037e366b9b2b8

Observation 61d387c8-950c-4507-82b2-d3eb05f6a283 · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Physics of language models: Part 3.1, knowledge storage and extraction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.531896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.531896Z digest=sha256:641adb544d2a476f1b6fa358abdeb1ea5c862d10afb48b3beffa81c25d44e5a8

Observation aae00ee6-8ee0-419b-adff-b23b52bf598b · outbound

This paper cites The urgency of interpretability, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The urgency of interpretability, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.538644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.538644Z digest=sha256:faa7e840b57a889ee05e597f82fc2a9c2a3c0ef020fb54d403e0b73f81474628

Observation deff4b33-b745-4e3a-9e7a-86e1815bb557 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.544723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.544723Z digest=sha256:2f71a73780357f7aa78c96d1180d45c1bea0037b2d1a5281fdcae7e51e469c2f

Observation 964da59b-029d-4588-8631-c9b228d56532 · outbound

This paper cites Ai as systems, not just models, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Ai as systems, not just models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.550537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.550537Z digest=sha256:45ef746642145332389263d2f9a8c5e77c951eefc0f28ae0a93fab0a791aba48

Observation 7c179344-ca3a-48d8-82d1-3ba93e274d41 · outbound

This paper cites Standard saes might be incoherent: A choosing problem & a “concise” solution.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Standard saes might be incoherent: A choosing problem & a “concise” solution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.556998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.556998Z digest=sha256:0b83858a83864b76f132ae5ffac9cbde7a45f3bb12d87d8af32580d7fe152ebd

Observation 1000c22c-0b9a-4e30-8e2f-592afc155ec8 · outbound

This paper cites Position: Interpretability is a bidirectional communication problem.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Position: Interpretability is a bidirectional communication problem

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.563027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.563027Z digest=sha256:d1d138e7721af9a4c824e0bcc2827b69a0785edd1776a69e017f16fdb6a6673e

Observation 237b7281-268c-41a9-b03e-66dee8e8639d · outbound

This paper cites A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.569755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.569755Z digest=sha256:d7db5d86eb30a07f8e1ba9b8e24b8089c2689f2e416bcae3f5b0acc0a33e30e3

Observation 2f798255-8dd9-4727-8730-208b6e6b64e5 · outbound

This paper cites Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.575857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.575857Z digest=sha256:607fe191a5d8e777b034b1e10150afe0a384d7a520041164b1a500ce5e0f4e0c

Observation e326204e-4a5b-4139-967e-9d642129ebdf · outbound

This paper cites Novum Organum.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Novum Organum

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.581703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.581703Z digest=sha256:774d314e582188ec5f535f7f5f16fe0132de61c9b536b7220d85fb5a82c74ed2

Observation 221e6ee1-3828-4e38-ba66-1b3543b46e24 · outbound

This paper cites Simplicity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Simplicity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.586827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.586827Z digest=sha256:d847358dfc2e8e81973be05b3b8c0f1ffa1830571ae2af8e56f90e3d2e391b78

Observation becd414b-4fff-4ff3-954f-598d5526b840 · outbound

This paper cites Design Rules: The Power of Modularity Volume 1.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Design Rules: The Power of Modularity Volume 1

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.592337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.592337Z digest=sha256:c5f65c24906be683f8ae8b49b73d7c9857b23f34909129b4481813a59f14ffdd

Observation 98a5edf3-3bec-4513-b1e0-5c1ecb7f32cc · outbound

This paper cites Icml 2024 mechanistic interpretability workshop, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Icml 2024 mechanistic interpretability workshop, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.597081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.597081Z digest=sha256:0d842f26a48182a6c3507378025206c537df10e69f937b092f3e76b3aaff3dbc

Observation 64e4beca-6ae9-4858-b525-195cbebf0549 · outbound

This paper cites Local vs. Global Interpretability: A Computational Complexity Perspective.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Local vs. Global Interpretability: A Computational Complexity Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.601776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.601776Z digest=sha256:6ae6870ed67d8a55c86df986364b6f68f4312c5a229d7d11f105b9f509032e67

Observation 3fa773a3-d922-4b64-81cd-e726fb950bf8 · outbound

This paper cites Explanation: A mechanist alternative.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Explanation: A mechanist alternative

Reference 16

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.262246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.606747Z digest=sha256:400b83238902ac9c7331df7cd8bbd19cc016471c9af86a45234b05610d89207d

Observation 3c2e7bee-d2fd-4024-bc7c-1cb2370aa993 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.612261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.612261Z digest=sha256:67f6e455f9fe11a44fcb20e370d1295fb139a4cb5d3af5f8dc6f95d5d78b1eb9

Observation c90eb841-85a5-44f5-b15d-ee4d768cc0c8 · outbound

This paper cites International AI Safety Report.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii International AI Safety Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.617272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.617272Z digest=sha256:2a5e9fe9a136b20c82f06c1a00867d56ad32a5f66c00fd2b3030651059430a86

Observation ec496079-ff91-465c-b9db-71c5d491d4cc · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic Interpretability for AI Safety -- A Review

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.622447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.622447Z digest=sha256:6e00663e963cac7591e16cc3b5080730be6485f5a300622aff57d2faee0dfa70

Observation 03f4349d-828d-465b-9065-1c382d45670d · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.239659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.627697Z digest=sha256:bcc26c0684611d0b592dff5e75808deecf7897616a2ce839b98fb49df7d00b00

Observation c960ebea-2022-4f48-a643-0b57db62aa68 · outbound

This paper cites Auditing local explanations is hard.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Auditing local explanations is hard

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.632693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.632693Z digest=sha256:01a7637ac337a9c8dd35e76c502b1deff83764ece030c9469c244ead11b7d04b

Observation ed2767c7-3100-4b87-beb9-87dce637ba3a · outbound

This paper cites Language models can explain neurons in language models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Language models can explain neurons in language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.637441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.637441Z digest=sha256:85b6262015370668e268450353e6176414037347c62d4ec609bcffb8f2dfda0a

Observation 87cc22af-9945-4683-b20e-ef73fe38ec6a · outbound

This paper cites An Interpretability Illusion for BERT.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii An Interpretability Illusion for BERT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.642287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.642287Z digest=sha256:3b5ea6d01e942d66f7dab8ca80f6be19e234b637cd9ab5e538afae3e00ab26e9

Observation a2b41a39-4f83-4cfd-a9f0-3663c232c3da · outbound

This paper cites Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.647126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.647126Z digest=sha256:a3c2979f8a2307c10a5263e61378e0cb978bf18924c406bae85478386b2c10f3

Observation 3a76e277-6588-436b-b83d-5be24a5b09ec · outbound

This paper cites Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.652281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.652281Z digest=sha256:445f8ac2c75219c805301bd3dcbb4301776029a38de68132260b01c045730962

Observation 135608c9-8774-44fb-bf3e-18692e656f90 · outbound

This paper cites Towards Monosemanticity : Decomposing Language Models With Dictionary Learning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards Monosemanticity : Decomposing Language Models With Dictionary Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.658022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.658022Z digest=sha256:93ccc627bdf5e25b8836fc85781a4884c634c7edbeaf0e11c5430b8074aaacd9

Observation 34747481-f21c-4452-9e82-01aac17a063e · outbound

This paper cites Propositional Interpretability in Artificial Intelligence.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Propositional Interpretability in Artificial Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.663766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.663766Z digest=sha256:f2e097f4b8bbd3c5781cdbac3f3d2fd02012424d2ae99ac893e33d197e46b6fc

Observation 8c79eff4-85ec-4984-8618-b2188376c0c3 · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A toy model of universality: Reverse engineering how networks learn group operations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.669143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.669143Z digest=sha256:88b6033da4a3c9c49de420fb4be96e61a80f1f756c3390d38e944947dcf0f32c

Observation ab6e7c15-147c-49e3-a67e-2ee2331be54d · outbound

This paper cites The evolutionary origins of modularity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The evolutionary origins of modularity

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.674171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.674171Z digest=sha256:ffacffdd402bb330347e24d9780181d8dcb01b074d088940f8fc84aabde39e4a

Observation 0bbc4ffc-89aa-4a8a-9228-f21bfc00c894 · outbound

This paper cites Towards Automated Circuit Discovery for Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards Automated Circuit Discovery for Mechanistic Interpretability

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.678871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.678871Z digest=sha256:a5e0f44c54329c052aaf7886485390ebaaddff3ee99eab2de77298ea5f0bfef2

Observation 288ced46-da29-40db-ae02-4244fd08479c · outbound

This paper cites Central dogma of molecular biology.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Central dogma of molecular biology

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.684525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.684525Z digest=sha256:5feda80ece456a97a20d64950593558324b08213c4ec74c9b3023b6901703939

Observation 0076b1cc-574b-4515-bbd6-8b68134604ec · outbound

This paper cites The beginning of infinity: Explanations that transform the world.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The beginning of infinity: Explanations that transform the world

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.690760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.690760Z digest=sha256:46f01f45c0feb771647da1ade591131b4491c140992bdc74a6129a58ae3daccc

Observation 9e2e0718-96c4-4dee-82a6-24b4caa1f973 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.696683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.696683Z digest=sha256:e959bb0f63e8694d71d77bd1db082e95ea56626d9d0a9aa7d0c77a45319d95fb

Observation 1643f022-9fd7-4021-af3f-81f37f4af433 · outbound

This paper cites Einstein.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Einstein

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.702151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.702151Z digest=sha256:d2e59ae07d8526aa0b1bf41f4be0d51eff61cd599591fb76a02f05d8aec009d7

Observation 6badf7f6-0bd7-48dd-b6d1-dbc160e7f37b · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.707415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.707415Z digest=sha256:3fb68bdc15225374af45963713daf75e30c07b420b85701431f6fbd33e9c1ada

Observation 465f48a9-16d1-4854-9556-2b1edce37af5 · outbound

This paper cites Clusterability in Neural Networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Clusterability in Neural Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.712412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.712412Z digest=sha256:a0dbc0449a5362aed7e7bc9876d28c2d55ec2aac9c8a1226294a46bb5820dc21

Observation d72172f6-6ee7-4cf2-a87e-c8aa458a6a2e · outbound

This paper cites Interpretability illusions in the generalization of simplified models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability illusions in the generalization of simplified models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.718000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.718000Z digest=sha256:85c187fa769d38503ac28bccef2c6a0adec9405210f3748ce1243ee5fca69a27

Observation f97a349d-416e-4ec6-ba95-63ed4aa9bd7f · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Scaling and evaluating sparse autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.723010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.723010Z digest=sha256:1a7580796f858eaaf222f77580edd868702b25407d95940bfa15ba6460e1d94e

Observation 712695dc-797c-4577-a954-6c0f803ec29b · outbound

This paper cites Causal abstractions of neural networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal abstractions of neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.728174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.728174Z digest=sha256:d93c504d05379bb430ee46b4514ded8ea04093047dc5fbed765858794c120ef9

Observation 22aea982-0987-44f7-ab5b-592ababe55dd · outbound

This paper cites Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.733991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.733991Z digest=sha256:ca170fde866d3490b2e236206f7807850a98dc38378e8b351695184b49d64fe8

Observation 25eafe53-c2b4-44c6-a494-fd448e1bf91c · outbound

This paper cites Clustering algorithms.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Clustering algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.740650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.740650Z digest=sha256:404aa653f060f850eeaa2484346c3812e0fe1dff04c6c00ef3b79adfad495f6a

Observation 5d25dd86-a4ae-4bdc-bad7-135601377bcb · outbound

This paper cites Compact proofs of model performance via mechanistic interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Compact proofs of model performance via mechanistic interpretability

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.746673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.746673Z digest=sha256:d44274048e436cad4be3a6d1789b69cb22a6b404a5200918d4738d7af3686ac4

Observation 33317061-ea4a-44d4-a1cc-f3220fce6bba · outbound

This paper cites Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.875795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.752590Z digest=sha256:76e5c17b4a51f1c855f10003fe549323dd6fa7a284b8925c5249ce342e1aacf8

Observation 6c1fbada-0a31-460d-8864-8ab9e751bc15 · outbound

This paper cites Hastie, R.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hastie, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.856843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.757293Z digest=sha256:71b3b97ceb123efeadf03fadd55f140a3c51a7a69a8afbeadcdc265c40f49596

Observation 6bb4f5b1-e9b0-4de3-b18b-8a9922be0e2e · outbound

This paper cites Hempel and Paul Oppenheim.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hempel and Paul Oppenheim

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.837001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.762261Z digest=sha256:64751099394f7c20ae809823a6ef5e34c6babacad2d2c82c1ea96e1162065a61

Observation 0fc0361d-3d60-41e5-9e5c-c8186b9233c4 · outbound

This paper cites Philosophy of Natural Science.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Philosophy of Natural Science

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.816805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.767579Z digest=sha256:7f37581cc37e19a4ff6b950452eecd862537b5d642a953936e7b6479ee4fac1d

Observation e355e61b-deec-4081-8248-fe70d5984b8b · outbound

This paper cites Bayesianism and inference to the best explanation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Bayesianism and inference to the best explanation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.797867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.773005Z digest=sha256:ee1f205ca53d4b6bc054939dfba295e52987f67e2026fabb8dd85319ae623b60

Observation b0510d13-33cf-4cb9-bfb5-b1bc7a98b276 · outbound

This paper cites Loss Landscape Degeneracy and Stagewise Development in Transformers.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Loss Landscape Degeneracy and Stagewise Development in Transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.779468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.779468Z digest=sha256:5273493fa210666e8a7fff44a7313c3b87f6515fb36e4e84925388a2a890fca7

Observation c5ec30f4-92f1-40f8-9acf-a47b1637be70 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse autoencoders find highly interpretable features in language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.785160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.785160Z digest=sha256:66e138d2471746c2f9500f06e94573923c081d289953722c3082d01cca10cc22

Observation 40da69cb-7913-46ec-9031-94e2f0f81322 · outbound

This paper cites Hutter, E.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hutter, E

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.767475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.791367Z digest=sha256:c70d271a8993277ee94b69db046ae573fc5c5235caf91923990ca237c0c33344

Observation 1c3adc74-c6de-4fdf-b320-a17fec806d6d · outbound

This paper cites Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.749269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.796480Z digest=sha256:970d639d44efe863b2bb4c99fc6468a1ba3e91b691b50a80eaf21d0fe007fcf7

Observation 7ff9effe-ae4b-46a8-be2e-3ce3887cac71 · outbound

This paper cites Kandel, J.H.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kandel, J.H

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.729487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.801635Z digest=sha256:970e57ee85991cc04877c25f8ccbafa95eb0d7e0de828523f2ef35a488b6cdef

Observation 6d585e48-eb4a-4bc0-9aac-d340d76ab5f0 · outbound

This paper cites Saebench: a comprehensive benchmark for sparse autoencoders, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Saebench: a comprehensive benchmark for sparse autoencoders, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.705111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.807554Z digest=sha256:b006f5e98af883773b8297607d685d6176b49e2083266d16f358c31997ddd7a6

Observation fd84da5c-f9e0-43b8-96b3-ed972dd73d61 · outbound

This paper cites Kennefick.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kennefick

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.687071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.814962Z digest=sha256:366ac1860aff3ffbcea8cf891418560a4bc185e6d0751778ca82deb43982ee27

Observation 16b197c7-a52f-41ad-bccc-66f48285f750 · outbound

This paper cites Explanatory unification.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Explanatory unification

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.667802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.820191Z digest=sha256:4203be646db8eb75819510a6a5408002df7dc244f8a395f7cd322a496cfad4fd

Observation 10e71510-9e4a-4a0e-ad3c-0aaf7a4baba4 · outbound

This paper cites Three approaches to the quantitative definition ofinformation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Three approaches to the quantitative definition ofinformation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.646435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.825015Z digest=sha256:b5be66a8d81039c786e92d1ae5a0159e32fa9ea7f1dcb06211b736b772f907e6

Observation 17746ef1-137a-4c4f-b7fd-8d3fa293643b · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:25:07.627680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.830341Z digest=sha256:0d08650279770369860cbe676f9ca9350f2037553f0a0ccee3ef43c51a3293ca

Observation 1587381a-db0b-4c29-8c54-0a0c152ca444 · outbound

This paper cites The Structure of Scientific Revolutions.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Structure of Scientific Revolutions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.607038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.835464Z digest=sha256:8c965ab5a3567c708dc154d513147237e1233bc044e861a9f7e171bf2972e04b

Observation aaa33f02-0ff5-4a53-aada-a5d9be8f6d4f · outbound

This paper cites Falsification and the methodology of scientific research programmes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Falsification and the methodology of scientific research programmes

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.586373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.840481Z digest=sha256:1f94a35673fe7db954bb3a67404e0c6b0895b90f3d80b571b186d8184e8b6f58

Observation 91c964fc-0773-420c-8279-fc1eb53403b3 · outbound

This paper cites The Methodology of Scientific Research Programmes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Methodology of Scientific Research Programmes

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.567371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.845034Z digest=sha256:502a1310e13c302fe9ddc7653dc85cab289b1785582393c5024f7f979cb56062

Observation f9885b31-5dd2-4dae-97f1-f880fb428f8b · outbound

This paper cites Sparse Autoencoders Do Not Find Canonical Units of Analysis.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse Autoencoders Do Not Find Canonical Units of Analysis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.850864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.850864Z digest=sha256:93bf982e0a314aed64f85d887e18955ad44f90f7267956e952d0b9ca7d64a155

Observation f7a9f288-c896-4fba-9c44-d8753e20975f · outbound

This paper cites Lindsay and David Bau.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Lindsay and David Bau

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.856242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.856242Z digest=sha256:ec024184858ba2c7dfc26695eb28cceec3b9564c0ca7db55948734ff957f21f2

Observation b0ae1407-4a60-447c-8d2e-dada1f82ccfc · outbound

This paper cites The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.862843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.862843Z digest=sha256:20eaa44f98459b619cb564268c6576dd942f0fe896ed0a9ba743cec65551cab5

Observation 9eb3975f-8e1f-4e66-a474-ec9c3b1a62af · outbound

This paper cites Mechanistic mode connectivity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic mode connectivity

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.539180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.868951Z digest=sha256:94c6f10ad17542c6683cb6a955de56c6a42ab480dd31d688ce1a23511f9c8230

Observation ca10ed64-93ed-4ac4-a197-12840824058c · outbound

This paper cites Information theory, inference and learning algorithms.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Information theory, inference and learning algorithms

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.874084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.874084Z digest=sha256:cdd68107a1568086ba383b946fdea97569f8b3cb8399a1d15a58cf87eb63a90f

Observation 86793f78-8c9c-49a5-b6b4-0c5e352fa2df · outbound

This paper cites Is this the subspace you are looking for? an interpretability illusion for subspace activation patching.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Is this the subspace you are looking for? an interpretability illusion for subspace activation patching

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.507881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.879925Z digest=sha256:649e3e5e543589578dc6812ae4f864b842de9d59e08e98b5adb485ace06bad2a

Observation 2085c2d2-ad9b-49d1-9c1b-412de83b2a98 · outbound

This paper cites Downstream applications as validation of interpretability progress, March 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Downstream applications as validation of interpretability progress, March 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.486729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.885763Z digest=sha256:619c9e6a842cd95d939e420b7c3af952709c040d44ab8290e4283fb558e4111e

Observation 484ea65f-aec7-463d-80ea-4d5c9dc4fd29 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.891673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.891673Z digest=sha256:a9133bceca3e04d2343ff81a1f6aa7cfd87af364f8d28f8e0fc4bd5339582118

Observation eb401795-6468-4d96-8e49-cfb9a832798b · outbound

This paper cites Cognitive styles in two cognitive sciences.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Cognitive styles in two cognitive sciences

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.468750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.898471Z digest=sha256:7c0627ed6e62d64ac4106fd45b3c8112f96280ad83679dedf04c8f1cbef5dbc4

Observation fb3e9be1-3e90-4402-bf7a-be0ded544915 · outbound

This paper cites Zoom in: An introduction to circuits.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Zoom in: An introduction to circuits

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.905678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.905678Z digest=sha256:2cbe7984ca3f2981349bbdb2f8e0c5f8a63c4f3c82e79d4932d3f143eff77411

Observation 536f7abc-d582-4db8-9057-fae24a5dba45 · outbound

This paper cites In-context Learning and Induction Heads.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii In-context Learning and Induction Heads

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.911761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.911761Z digest=sha256:31bcabbd76133ec7b0abfd7edaa73dad09dac5cf59a237e85a8031474aea64b3

Observation 293d1e50-1950-4e08-84f3-51ad6d5bb269 · outbound

This paper cites Automatically Interpreting Millions of Features in Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Automatically Interpreting Millions of Features in Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.918441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.918441Z digest=sha256:69d186807c47028fa2183bbfd4142216da2d59737524bc56743a6ff6ebe3d453

Observation 3b0b26c7-9cc4-4b96-8dd8-8f02df58426d · outbound

This paper cites Causality.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causality

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.925700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.925700Z digest=sha256:6299a2e52d40ad9abe7388584267248546082640778441b9b28e15991cd60a6b

Observation c6aebdff-6c3d-404e-b0fd-1062e6f53578 · outbound

This paper cites Poincar \'e.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Poincar \'e

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.416845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.932748Z digest=sha256:055205eed8a0810fbe84ff4abcea4181f009dc9897761c7824050fe601dd2de7

Observation e2af5c31-1409-4f9f-b0c2-e63b82fc5688 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.944080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.944080Z digest=sha256:9dbbaa98efed3fe4c137feba421b5945f5ad09199cc8104ecbec5607e82a7388

Observation 5b60c837-b55c-4963-8912-cee876b825cf · outbound

This paper cites Hume on theoretical simplicity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hume on theoretical simplicity

Reference 76

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.221515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.951779Z digest=sha256:1c7dbe80d001eca7348acd483d560856feccea6565e6e856d66be4deb37adef4

Observation e7de4f8c-c5f8-4c6e-bdb3-6d566e58bd11 · outbound

This paper cites Escalation risks from language models in military and diplomatic decision-making.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Escalation risks from language models in military and diplomatic decision-making

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.962513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.962513Z digest=sha256:eb60515a88293ec5f232904b7e469f68724a357681b0085149974a7addda5094

Observation 4747c415-0ae3-4b04-b1ee-2029e9bd4a1f · outbound

This paper cites Four decades of scientific explanation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Four decades of scientific explanation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.384390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.968237Z digest=sha256:bde09a220039cf5d924cabfb93ca4b17f8b370c69bce8c6bc3e64aa38eb6b261

Observation 80943169-dfbc-466b-a923-efe22d9afad1 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:25:07.365849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.973050Z digest=sha256:6155324ba47a54f1f93ef0f4aa651734a83cbb5463986be400868093ebfd38d0

Observation 29c5b0f6-51c4-4735-99a3-2c9f05b2f048 · outbound

This paper cites Mechanistic?.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.978424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.978424Z digest=sha256:6493086c5ba997339680160ed3ef2eac159fa07c8e58b827663e25742f9b8249

Observation 08075a01-c91b-4cc8-8e58-12d0004ac73c · outbound

This paper cites Theoretical Virtues in Science: Uncovering Reality Through Theory.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Theoretical Virtues in Science: Uncovering Reality Through Theory

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.346087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.984022Z digest=sha256:7a4e5835517b3ffbfb956ab5b7dc1fd3690e0f72dc1982eb2cfef27e05d7a90c

Observation 740ce7be-8b7a-4b84-b752-dd27ef2f759c · outbound

This paper cites Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.990103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.990103Z digest=sha256:6309f11a716b6a832365af4da62ebb00e4c57c37f41101372dc76d3633c7bdf7

Observation 1b7789d5-9c8b-4ace-9512-ac09ab787734 · outbound

This paper cites A mathematical theory of communication.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A mathematical theory of communication

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.996982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.996982Z digest=sha256:7a518937a8b116b94243e0607446687266a1588d1b24f1fc4e58d958701dd5b6

Observation b417d3b7-a3fc-4a92-a138-97465a1fe086 · outbound

This paper cites Open Problems in Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Open Problems in Mechanistic Interpretability

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.003123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.003123Z digest=sha256:68193045e2775732fafc500a40ababce3feb5beac63474b86f0aa2eaa0811ebf

Observation e07c0ebc-f326-427a-a03f-bbc9cbaf8612 · outbound

This paper cites Hypothesis testing the circuit hypothesis in LLM s.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hypothesis testing the circuit hypothesis in LLM s

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.296737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.010034Z digest=sha256:d83742bffb80ea7fe3719aa411988cc1767337ab6687ca4b40b2c2bdbf4a8395

Observation 1d512897-3ac1-4f50-b804-40995ead7c76 · outbound

This paper cites The golden mean of scientific virtues, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The golden mean of scientific virtues, 2024

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.276920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.016375Z digest=sha256:7ccd5e71315efe4b600790cc954ee3a158fe37f669ebf1cd0c22d2d03891ed98

Observation 1d7ce50c-4f21-4f8c-a2d0-82bfdc5f1ff6 · outbound

This paper cites Knowledge in Perspective: Selected Essays in Epistemology.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Knowledge in Perspective: Selected Essays in Epistemology

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.259403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.024761Z digest=sha256:b93eb39e5c4e94df6b5fbd8b32ed57296245a2757ae3c03f65deb49df9d176f7

Observation 352006f4-c9b9-4bf7-bc3c-71372f66f6b8 · outbound

This paper cites Grokking group multiplication with cosets.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Grokking group multiplication with cosets

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.240232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.031559Z digest=sha256:f4a895ef610009d3e68e1d24372cd1187f4b1dc78f230f38d9cb7dedf00ff70a

Observation ca0eab37-098d-4e05-a20d-024b68c63bb8 · outbound

This paper cites Simplicity as Evidence of Truth.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Simplicity as Evidence of Truth

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.218842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.037028Z digest=sha256:dbb40201edd3d4228d743c2ec4f400559a709227daff1fdc4080e81021ba150c

Observation 2e2cb470-0087-4b83-a9c2-b26796256c1f · outbound

This paper cites TracrBench: Generating Interpretability Testbeds with Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii TracrBench: Generating Interpretability Testbeds with Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:25:06.330001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.043348Z digest=sha256:104ef251074987b232e6ce0d4d5568a9032fa1501257ddb872e2a70ed73ec3e0

Observation 0a2c7ecf-696a-4a1a-9927-dc96616081c7 · outbound

This paper cites Interpretability in the wild: a circuit for indirect object identification in gpt-2 small.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability in the wild: a circuit for indirect object identification in gpt-2 small

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.193035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.049199Z digest=sha256:b1564dc38ca3e114e8734cb110de549d09c3fb1ae4aed23a729ec38bbab995aa

Observation 57fadde9-0811-439f-84ad-fa1f01846649 · outbound

This paper cites Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005

Reference 92

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.202634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.057524Z digest=sha256:f73f34db997e9d72af58d53e3d28a270f9b9a49fa87a1406261df2b11767e7d5

Observation b37e3d23-7ad3-4fdd-b774-1a560c33b6d1 · outbound

This paper cites Understanding as compression.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Understanding as compression

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.173192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.064714Z digest=sha256:3045accbc63d1463fbbf4bd287101bd0140039b4df05a40fbfdbf323f2ecaadb

Observation 095fdfb3-67ab-48ae-a712-b6d6f312a35d · outbound

This paper cites Geschichte und Naturwissenschaft.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Geschichte und Naturwissenschaft

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.154724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.071370Z digest=sha256:710d934a8fd6fcf0bbd6bcdbaa9862618ce8590bdc7606d6ad70378cce7b69e3

Observation 721adab6-90b7-48eb-81f9-c03e9c9c1263 · outbound

This paper cites From probability to consilience: How explanatory values implement bayesian reasoning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii From probability to consilience: How explanatory values implement bayesian reasoning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.136110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.076238Z digest=sha256:30234f68b9451e3b61715ff0f18801d76648567014ccd3d04485df82551fb063

Observation 0c046019-3c7f-41e4-9991-f9a048c1a470 · outbound

This paper cites Woodward.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Woodward

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.111381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.081518Z digest=sha256:b3e9c14be19040c69a2ec77972c42cc86c069a26e6cdba5a113b142a2a1d01cc

Observation 3ff95413-4b01-4574-8af5-5fffe4e5c756 · outbound

This paper cites Towards a unified and verified understanding of group-operation networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards a unified and verified understanding of group-operation networks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.086534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.086534Z digest=sha256:6fd2197c8a290ae4bbe31f8bea16d14e238759e8f049071f4e2267eec98fb3e9

Observation 74a24061-0e17-4599-a309-ec652df7de8d · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.093358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.093358Z digest=sha256:229447b604c462ed0981f2175a730b1e400b7ce151a06df6bea53c803e7f167c

Observation 5268333d-c642-4bbb-b2eb-f957abac149a · outbound

This paper cites A theory of usable information under computational constraints.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A theory of usable information under computational constraints

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.101369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.101369Z digest=sha256:fe210f7ec12635caf8ee5ce33c06e082b9a51517c9ca6154d43716eac624bb7c

Observation 7474a404-cacd-4e6c-bd4e-726f49f34e19 · outbound

This paper cites Locally decodable codes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Locally decodable codes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.076567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.108473Z digest=sha256:757b14cc76a03e8cbc38188930c447e6b4d3cc470a129bd7a31cc5f52b699a96

Pith citing papers

Observation c7a68c25-a842-4499-bca8-0cea977ccc1b · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.444788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:98a737ab408ba623b4f6e61b84bae3291b5c566c6b12a95d117a65a77ca285bd