Pith. sign in

Paper Citation Record · LEDGER

MIB: A Mechanistic Interpretability Benchmark

As of 20 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 11 inbound Pith citation observations for arXiv:2504.13151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13151 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:44.485555Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:30.524426Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 674369d7-d84f-487f-beaa-02a27407f18e · outbound

This paper cites write newline.

MIB: A Mechanistic Interpretability Benchmark write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.094644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.094644Z digest=sha256:4cf7e931c964aac4b0b3f2f6005fb7b9801f6f3247641a195df8958d9ab633c6

Observation a5be7bd8-9474-4e1d-825c-1eab4f215b6a · outbound

This paper cites D., D'Oosterlinck, K., Feder, A., Gat, Y.

MIB: A Mechanistic Interpretability Benchmark D., D'Oosterlinck, K., Feder, A., Gat, Y

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.437930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.101547Z digest=sha256:5675e7d8a87c4e6b78a523809208e27f867829a9f9db14765054fe02ca684658

Observation a17374ba-4091-407d-9def-e6769eb67fcc · outbound

This paper cites Naturalistic causal probing for morpho-syntax.

MIB: A Mechanistic Interpretability Benchmark Naturalistic causal probing for morpho-syntax

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.106861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.106861Z digest=sha256:09ce3d84bc83837ab7d165bff97850a399783149e4b85b44f32671bc5cab7a59

Observation 2fa739a7-e82a-4afd-b394-598ccfca6746 · outbound

This paper cites Causal G ym: Benchmarking causal interpretability methods on linguistic tasks.

MIB: A Mechanistic Interpretability Benchmark Causal G ym: Benchmarking causal interpretability methods on linguistic tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.111475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.111475Z digest=sha256:cf8690c838152d0221218a4b41d3b7eb2e8a28cd0376195aa74495d93b4968b5

Observation 01cbd3b5-393a-461f-9159-7dc522c034b0 · outbound

This paper cites G., and Augenstein, I.

MIB: A Mechanistic Interpretability Benchmark G., and Augenstein, I

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.116361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.116361Z digest=sha256:007d77c013cf0fc6acc02913bc6c3dce953cc1bf4f2bb4e43bf45432a0e07c00

Observation 16fd34b1-ccb3-4864-9c57-c8b6369521ab · outbound

This paper cites E., Hume, T., Carter, S., Henighan, T., and Olah, C.

MIB: A Mechanistic Interpretability Benchmark E., Hume, T., Carter, S., Henighan, T., and Olah, C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.427497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.122079Z digest=sha256:4f3831948503bb4de60ba21bf707c3272bb25dc8cbc73a72040dbb30f26ca346

Observation 93ff46ae-35ab-47ac-846c-c5d852226093 · outbound

This paper cites an unresolved cited work.

MIB: A Mechanistic Interpretability Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.126664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.126664Z digest=sha256:59e736d270a4b3c34000b56d8eb8ae893c0cf66a0134def559c3c205756f0251

Observation b85ee3b9-aea6-423e-8014-aee498eda74f · outbound

This paper cites D., Schlichtkrull, M.

MIB: A Mechanistic Interpretability Benchmark D., Schlichtkrull, M

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.134591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.134591Z digest=sha256:4d539ce80e4cedf2f1ab9d4c9a7eb5a2277417187b52545d4adf17080e3c5b43

Observation cafa9268-e036-453f-ae28-89af066c3c09 · outbound

This paper cites D., Schmid, L., Hupkes, D., and Titov, I.

MIB: A Mechanistic Interpretability Benchmark D., Schmid, L., Hupkes, D., and Titov, I

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.140323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.140323Z digest=sha256:d32fef67c7e9adca9d9a9902db0ecbd3ff5089f0f799d35514148f03aa0744cd

Observation 95b50249-2a8a-4d33-bd88-82c9678d0523 · outbound

This paper cites Causal scrubbing, a method for rigorously testing interpretability hypotheses.

MIB: A Mechanistic Interpretability Benchmark Causal scrubbing, a method for rigorously testing interpretability hypotheses

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.146910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.146910Z digest=sha256:1e1199a700c20b4785d499c3af4b5eed432638df836151cde46a597f74e0a7a2

Observation ef87e1d5-8e18-4197-aa9f-d7899a7711c9 · outbound

This paper cites Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small.

MIB: A Mechanistic Interpretability Benchmark Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.152348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.152348Z digest=sha256:1a5edcb24de26ee49ce53e6e65095b44f8e066260cbfa1377d87b426074050fa

Observation 454e5ae9-d5f6-4a31-822d-38d018706c9d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

MIB: A Mechanistic Interpretability Benchmark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.157193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.157193Z digest=sha256:e3def7a032e9ab024cecd9f34266c171f75f3e67f195326dcc9a62900320d83a

Observation 025ca9df-c481-4d15-ba75-0fcad7a4efae · outbound

This paper cites Evaluating the ripple effects of knowledge editing in language models.

MIB: A Mechanistic Interpretability Benchmark Evaluating the ripple effects of knowledge editing in language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.161490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.161490Z digest=sha256:9a965f132a6663e6fcfc49e62e52ac02af424933b736bd20e4ada71a0575780b

Observation 3d0d4a95-dee1-4586-a5e2-38e9c9bb54c0 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

MIB: A Mechanistic Interpretability Benchmark Towards automated circuit discovery for mechanistic interpretability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.165647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.165647Z digest=sha256:a1b7f9fe5703d661998df38864411d3a66822efb934acaf5d78c871ab17bc3a8

Observation 4c891117-16f1-494d-83e8-14c797c37029 · outbound

This paper cites Are neural nets modular? I nspecting functional modularity through differentiable weight masks.

MIB: A Mechanistic Interpretability Benchmark Are neural nets modular? I nspecting functional modularity through differentiable weight masks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.396826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.169774Z digest=sha256:bf88fc7b9bfc289c3b47c42d99feae6f69da9662b9f3f5512ed8f985cac26d6a

Observation ab60f239-d8e0-4621-9aee-6d56a85db3bc · outbound

This paper cites D., and Geiger, A.

MIB: A Mechanistic Interpretability Benchmark D., and Geiger, A

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.385893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.174625Z digest=sha256:1cf54b98164b254574eecd258c684b4802de19275d89e15355fd14316cc4abd9

Observation 86eb13d8-49b5-4fa6-8a90-c4328d39b859 · outbound

This paper cites Representational analysis of binding in language models.

MIB: A Mechanistic Interpretability Benchmark Representational analysis of binding in language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.179106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.179106Z digest=sha256:a222d9a712113f09c14189393bac7d53d0093c4c2c5f48ee5a917c41ffa0efb0

Observation 9e775a01-b83c-4e15-a867-5396a83a81e1 · outbound

This paper cites Discovering Variable Binding Circuitry with Desiderata.

MIB: A Mechanistic Interpretability Benchmark Discovering Variable Binding Circuitry with Desiderata

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.183284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.183284Z digest=sha256:ccf84392495de8b9fa10177d19bb62ce1ae09f348612637378324eb82eb77ad1

Observation 81c66191-d82e-4b4c-bc30-4f7a22c119a5 · outbound

This paper cites The Llama 3 Herd of Models.

MIB: A Mechanistic Interpretability Benchmark The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.188043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.188043Z digest=sha256:891af803f1743735468d1045db6ee0b8803e74460f4de0e2157898675ee625d2

Observation 4ff8a6b4-950a-4360-89a8-922232fa8346 · outbound

This paper cites and Voita, E.

MIB: A Mechanistic Interpretability Benchmark and Voita, E

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.192154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.192154Z digest=sha256:7cf7e8aa03ffbcf9caee9adff36fca43f57287353ab65860f9a6b4523c936a1c

Observation 2cb92199-7342-4f6f-b00b-af85497cc4c4 · outbound

This paper cites A Primer on the Inner Workings of Transformer-based Language Models.

MIB: A Mechanistic Interpretability Benchmark A Primer on the Inner Workings of Transformer-based Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.196854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.196854Z digest=sha256:9bb64a3abbdad0a2dd53142dd454b183eb56f5f2dd457996ceabc6b81dce8296

Observation b767b015-fce1-45bd-a0a0-6e26202630f3 · outbound

This paper cites Causal analysis of syntactic agreement mechanisms in neural language models.

MIB: A Mechanistic Interpretability Benchmark Causal analysis of syntactic agreement mechanisms in neural language models

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-16T12:19:44.202047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.202047Z digest=sha256:4dc8d5b68d8465d1dd1a71e8f24d4cfcd5685bdcb0065d6f0bc13d7a9a03a75b

Observation 0e20172b-b8be-442d-a86a-924ebef1a8ec · outbound

This paper cites Neural natural language inference models partially embed theories of lexical entailment and negation.

MIB: A Mechanistic Interpretability Benchmark Neural natural language inference models partially embed theories of lexical entailment and negation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.206991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.206991Z digest=sha256:a7fab614d5f0ed3628e0b6d5e053fb3f76f85b43aa6a5a4b6cfa5bde47e4d9e8

Observation fcf78a62-d267-40a2-9454-b58de6ce52b1 · outbound

This paper cites Causal abstractions of neural networks.

MIB: A Mechanistic Interpretability Benchmark Causal abstractions of neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.211407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.211407Z digest=sha256:81953b07e53a2eeac0392cb1ead60dc1e462a366b32441ff6e5632c658d6d8cd

Observation 0d7cd494-b1de-4b9b-b4bd-199414ac0a7c · outbound

This paper cites Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability.

MIB: A Mechanistic Interpretability Benchmark Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.215403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.215403Z digest=sha256:03915471206e6f433039e1820d35b71a9c2916fa9193e3e00a000641ae1b20c0

Observation 159477ae-8199-4e6c-a43d-f6747a5a828f · outbound

This paper cites an unresolved cited work.

MIB: A Mechanistic Interpretability Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:19:45.367226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.221212Z digest=sha256:b1df289b8a5d2386468de30546c4dff3fc5b24f2244f5ad028748cc501f39940

Observation 67f1e680-6a0c-47d3-a9b6-b37b11327c18 · outbound

This paper cites Interp B ench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques.

MIB: A Mechanistic Interpretability Benchmark Interp B ench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.357306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.226590Z digest=sha256:eb3d4825928d8e12d54d26b7dabf683d57a99f92f2fb81c09ae56f805d211dd8

Observation e3a7e579-4ccd-442c-9b1a-1fe259cccea9 · outbound

This paper cites Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms.

MIB: A Mechanistic Interpretability Benchmark Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.347160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.231187Z digest=sha256:8bd078063dd6fc0787331798c1eb5cd52c2a5f589f5dc2ffa92bff57a4563658

Observation 3c8c3784-55e0-4e75-9d32-8e0e035ce8c9 · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

MIB: A Mechanistic Interpretability Benchmark Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.235129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.235129Z digest=sha256:dcc26ccc6b7c85573f570ac85745faaafdae3a3510b342892525362af12b30b6

Observation cd644efd-5788-419a-93ac-a7ec2fd61c91 · outbound

This paper cites We Can't Understand AI Using our Existing Vocabulary.

MIB: A Mechanistic Interpretability Benchmark We Can't Understand AI Using our Existing Vocabulary

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.240029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.240029Z digest=sha256:e89d511949a20aaf5b1b75f87535e5779acd43a8ecdd1f1c2f63467a23d3151e

Observation a7658cd3-5c92-46fe-9528-9509ba5e4241 · outbound

This paper cites RAVEL : Evaluating interpretability methods on disentangling language model representations.

MIB: A Mechanistic Interpretability Benchmark RAVEL : Evaluating interpretability methods on disentangling language model representations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.244630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.244630Z digest=sha256:26a437530dd3be1ddcb2910998d16df1b2f9619d3c9b0b1741ce4d26c4fc5e22

Observation dd3b7a43-7c8d-4fad-81b7-f560faacd2fa · outbound

This paper cites Unified view of grokking, double descent and emergent abilities: A comprehensive study on algorithm task.

MIB: A Mechanistic Interpretability Benchmark Unified view of grokking, double descent and emergent abilities: A comprehensive study on algorithm task

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.335766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.249555Z digest=sha256:daf2adcce6023dc21a8d7f70df666b0ff5d47f91f2716b3d86fabf10a619bff9

Observation ba885556-53cf-4d31-baae-1fa6f2581c70 · outbound

This paper cites R., Ewart, A., and Sharkey, L.

MIB: A Mechanistic Interpretability Benchmark R., Ewart, A., and Sharkey, L

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.253838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.253838Z digest=sha256:5e210f9f6fc8e497a6db72f21a3f1c97a158ac5be0bed0d01968b5824de642a6

Observation 2b843f47-cb30-4d96-8fc5-11855fac63ec · outbound

This paper cites Mistral 7B.

MIB: A Mechanistic Interpretability Benchmark Mistral 7B

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.259346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.259346Z digest=sha256:398b766039c777064f453e1daf075a06cf637f7a34097f96ac102f74e431c355

Observation 9c3ceb61-fde3-4479-8660-b8ac8ed54565 · outbound

This paper cites MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.

MIB: A Mechanistic Interpretability Benchmark MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.264200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.264200Z digest=sha256:104f9bf662d58dc5b7cb3a6fd9b5499292287870761e7c99b9be611165b8ae10

Observation 6e592365-1c9a-4636-b610-6edcf208eac5 · outbound

This paper cites S A E B ench: A comprehensive benchmark for sparse autoencoders, 2025.

MIB: A Mechanistic Interpretability Benchmark S A E B ench: A comprehensive benchmark for sparse autoencoders, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.316029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.269754Z digest=sha256:b5ed0516d7ae26f1dfe357dbd2389edc58e828bc16a6e771cb8c9776cf6862d1

Observation 3a193065-86e9-4816-bfa8-4f3f5cfcd559 · outbound

This paper cites and Janson, L.

MIB: A Mechanistic Interpretability Benchmark and Janson, L

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.302802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.273862Z digest=sha256:edfdeae9a46226b4e00fbd5c4054d3762737a251d0d726764563ec16b539313a

Observation c350eff6-16c9-414e-b9a9-9ddc1a92f23b · outbound

This paper cites Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions.

MIB: A Mechanistic Interpretability Benchmark Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.278545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.278545Z digest=sha256:1e74846348d3b9641e50daf4bf6114dee1d349e31845872ee5272a2d4892d5e9

Observation b56e95c5-7fb8-4489-902b-13c1ded5124a · outbound

This paper cites Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla.

MIB: A Mechanistic Interpretability Benchmark Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.288268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.288268Z digest=sha256:5260fc8b6c1f2a032a79fbd0fedcee9c5cd3ac0eea32513b1576468386057980

Observation c21ef519-b60e-4c89-a65e-d15845d80eee · outbound

This paper cites Gemma S cope: Open sparse autoencoders everywhere all at once on G emma 2.

MIB: A Mechanistic Interpretability Benchmark Gemma S cope: Open sparse autoencoders everywhere all at once on G emma 2

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.292354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.292354Z digest=sha256:b9b57aa0b176729570fcd90577d758d2fa8d247a30484bcdb812f885bf81d577

Observation 2910aa50-e957-464a-a09e-b15eb0c1cf74 · outbound

This paper cites J., and Tegmark, M.

MIB: A Mechanistic Interpretability Benchmark J., and Tegmark, M

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.291289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.296230Z digest=sha256:b4b7dbdde1198a4d6b8c27750cbd104ad3f6f85f3ce692615ad14949ddf47c12

Observation 0baac6ee-5c28-4dd0-8abf-5d619f2577de · outbound

This paper cites and Tegmark, M.

MIB: A Mechanistic Interpretability Benchmark and Tegmark, M

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.301616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.301616Z digest=sha256:5941ede4c44db88a89c86a41310f4f55cf303a8ad16f51783bf5d0cb958611e9

Observation e365039e-a2d8-4544-b83e-afe3d88d5f50 · outbound

This paper cites J., Belinkov, Y., Bau, D., and Mueller, A.

MIB: A Mechanistic Interpretability Benchmark J., Belinkov, Y., Bau, D., and Mueller, A

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.305937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.305937Z digest=sha256:945cfece5f44695e9c24ff262763e3c57d7ccd9c6893da9c834c2ad2c5cca56f

Observation 5f9028d7-99f1-403d-b8c0-05f8ddd9c967 · outbound

This paper cites J., and Belinkov, Y.

MIB: A Mechanistic Interpretability Benchmark J., and Belinkov, Y

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.310334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.310334Z digest=sha256:a8171c18d677a895d88d31ff2ea2d1508de21bb2d7938a53f729eb6e9feaf324

Observation 8347a5e6-4b61-4ee7-a6ec-8d22c4aadcef · outbound

This paper cites Circuit component reuse across tasks in transformer language models.

MIB: A Mechanistic Interpretability Benchmark Circuit component reuse across tasks in transformer language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.257719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.314615Z digest=sha256:c17a082bfbe7889c7c0ee09eb1101289efbbe91079a79ffcccbd7cc2924532ad

Observation 88b197ba-be37-447b-a5c9-253da501e000 · outbound

This paper cites Transformer circuit evaluation metrics are not robust.

MIB: A Mechanistic Interpretability Benchmark Transformer circuit evaluation metrics are not robust

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.246129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.318591Z digest=sha256:365a6632dbc6aff87a1c1ec04f0e156c1e2a1a086df5da0cd3fae9e3e5d40d75

Observation fa9aeabb-d5a5-4bfd-b2f6-26e127fc5b48 · outbound

This paper cites ALMANACS: A Simulatability Benchmark for Language Model Explainability.

MIB: A Mechanistic Interpretability Benchmark ALMANACS: A Simulatability Benchmark for Language Model Explainability

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:19:44.887552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.322707Z digest=sha256:5ca93de0aadffebefb3f1f8ad4f7ad29dff7e35131ad5e8571b169aff7a01a8e

Observation c055a362-2660-4f5e-bb09-28a5e43bac35 · outbound

This paper cites Missed causes and ambiguous effects: Counterfactuals pose challenges for interpreting neural networks.

MIB: A Mechanistic Interpretability Benchmark Missed causes and ambiguous effects: Counterfactuals pose challenges for interpreting neural networks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.235059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.327378Z digest=sha256:e369ee4cb9b13111f337773f7c7915d856285ce4a228adbf69e1db9e63621ced

Observation bfbcb86d-4514-4b96-be4a-6b61d1c914f1 · outbound

This paper cites S., Sun, J., Todd, E., Bau, D., and Belinkov, Y.

MIB: A Mechanistic Interpretability Benchmark S., Sun, J., Todd, E., Bau, D., and Belinkov, Y

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.332367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.332367Z digest=sha256:22294b80de1bf6436301e6cb2b9fec11e99167fe8829b570e840286d93723c8c

Observation cabedd9d-baae-48f7-bd64-1f0fe3171c7d · outbound

This paper cites Attribution Patching : Activation Patching At Industrial Scale , 2023.

MIB: A Mechanistic Interpretability Benchmark Attribution Patching : Activation Patching At Industrial Scale , 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.222449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.338034Z digest=sha256:d60229f11efb50542d76bd74aaf7a276f3247ea747971aa9f3811b887a9513e6

Observation d82fae5a-ff37-4040-a466-be814abc645e · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

MIB: A Mechanistic Interpretability Benchmark Progress measures for grokking via mechanistic interpretability

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.342604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.342604Z digest=sha256:2a281e213372c0958aab71deca3f039a7757c3a21e9aaa394176280fa83be321

Observation af003af9-d57d-46c9-ae48-ebeee85d7c9d · outbound

This paper cites Arithmetic without algorithms: Language models solve math with a bag of heuristics.

MIB: A Mechanistic Interpretability Benchmark Arithmetic without algorithms: Language models solve math with a bag of heuristics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.204428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.346598Z digest=sha256:a209dbbeef73e0bdf360362f18dd410bed4d4913f20e80f1e41f7c401218fa4d

Observation f7c2436a-759a-4d3d-93fc-2f73b73764ec · outbound

This paper cites an unresolved cited work.

MIB: A Mechanistic Interpretability Benchmark Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-16T12:19:44.575617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.351265Z digest=sha256:d952aa85db3c0e696e2d67cbe69ff5b802dd72e0316b80500bdb80970315137d

Observation 8a280f95-b32d-42ee-8efe-7172c061a9d1 · outbound

This paper cites Zoom in: An introduction to circuits.

MIB: A Mechanistic Interpretability Benchmark Zoom in: An introduction to circuits

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.355595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.355595Z digest=sha256:8f253ed1329ff76c01a9797606a17aee98edc558256b58a269d50fa67cd35077

Observation afb3294d-a223-47c9-95fe-e7e1ccb215ea · outbound

This paper cites Chat GPT.

MIB: A Mechanistic Interpretability Benchmark Chat GPT

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.193351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.360286Z digest=sha256:c0f67d11be1577740bec6a49f35b764fb3c4ee8b94a530da8f4e25f1a1837525

Observation e467ff9b-46bf-45ea-810c-7a1577522a47 · outbound

This paper cites T he W orld of an O ctopus: H ow R eporting B ias I nfluences a L anguage M odel`s P erception of C olor.

MIB: A Mechanistic Interpretability Benchmark T he W orld of an O ctopus: H ow R eporting B ias I nfluences a L anguage M odel`s P erception of C olor

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.365711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.365711Z digest=sha256:480bb8bf850144aa26e245b19cc461afb00e02c6a7bea24738bb4bf1000a61d4

Observation 617d4acc-1f71-43d3-a42e-bf911c5566b2 · outbound

This paper cites Direct and indirect effects.

MIB: A Mechanistic Interpretability Benchmark Direct and indirect effects

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.370867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.370867Z digest=sha256:2fcfc37c542e99ebc5e721456e4cc6c3e770921a0356ac20e3e26b5770ae1cce

Observation bc420875-3adf-4312-85f4-9b7e819f0da7 · outbound

This paper cites R., Haklay, T., Belinkov, Y., and Bau, D.

MIB: A Mechanistic Interpretability Benchmark R., Haklay, T., Belinkov, Y., and Bau, D

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.375601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.375601Z digest=sha256:284a606eda7f4f02e8cf21fddf19c286fa2a6fb223d2b0f4d3ddf9f8e6e97b7c

Observation 0c5e45ca-ee17-459f-9526-0288e992fc1d · outbound

This paper cites Language models are unsupervised multitask learners.

MIB: A Mechanistic Interpretability Benchmark Language models are unsupervised multitask learners

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.166129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.381041Z digest=sha256:5ce02dfd6119538b42cc77d1ba7e9dc28d051a3124a7647c5c9769f0796d0fec

Observation 7cb43e06-e96b-4033-955b-29f703262ae3 · outbound

This paper cites Toward transparent AI: A survey on interpreting the inner structures of deep neural networks.

MIB: A Mechanistic Interpretability Benchmark Toward transparent AI: A survey on interpreting the inner structures of deep neural networks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.385645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.385645Z digest=sha256:d989ce68a0af78aea26aeb228fe6f0dc6dfce2ff907694c354f1209ce25b0a3f

Observation 6cfe65e8-a641-4b35-a9be-8708ea59b0b0 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

MIB: A Mechanistic Interpretability Benchmark Gemma 2: Improving Open Language Models at a Practical Size

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.389775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.389775Z digest=sha256:6ba7334f8967d6619e6241dfbd8685ffc9d1f5db00ee0b50dd7d918e5452595b

Observation 8af88365-fba6-440a-8db8-8c2c9d5d63be · outbound

This paper cites and Wiegreffe, S.

MIB: A Mechanistic Interpretability Benchmark and Wiegreffe, S

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.394397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.394397Z digest=sha256:1cbe79a9e4d415c45cd2ce9ea887c6bc0b7a3833d4af1655dd4889f6b9187f48

Observation b335c0b4-7620-42a0-a07d-a8875fc54da3 · outbound

This paper cites R., Materzynska, J., Chowdhury, N., Li, S., Andreas, J., Bau, D., and Torralba, A.

MIB: A Mechanistic Interpretability Benchmark R., Materzynska, J., Chowdhury, N., Li, S., Andreas, J., Bau, D., and Torralba, A

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.144937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.398080Z digest=sha256:3f0e3b8afad1557076507d8126dd09cfbbc0be27a24580b17cf685e6c4573cb2

Observation bafc4f95-d018-43e8-b457-ab7e7916768d · outbound

This paper cites Open Problems in Mechanistic Interpretability.

MIB: A Mechanistic Interpretability Benchmark Open Problems in Mechanistic Interpretability

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.403261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.403261Z digest=sha256:8fe0911363f4fae9a66b61d04e6423db60046ce4896eb1d1576840d3be8200fa

Observation 2024cd65-c2f0-42b4-9238-235e5decdc44 · outbound

This paper cites Hypothesis testing the circuit hypothesis in LLM s.

MIB: A Mechanistic Interpretability Benchmark Hypothesis testing the circuit hypothesis in LLM s

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.131791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.408434Z digest=sha256:0e8064c9df78e3ee2d508ac34dfe956e19c8452c2511a0c9324e61ddd74afc37

Observation 4be3bb3c-596e-41ff-8f4c-73844cad5954 · outbound

This paper cites Neural and conceptual interpretation of PDP models.

MIB: A Mechanistic Interpretability Benchmark Neural and conceptual interpretation of PDP models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.120559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.413534Z digest=sha256:80395b7691004a9db8bcd6fcc81f62c4cddb176fc34aaeb5e94c4110b276b826

Observation 8b01a9d3-50db-44e1-9ef4-df33b7827098 · outbound

This paper cites A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis.

MIB: A Mechanistic Interpretability Benchmark A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.107760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.417522Z digest=sha256:154a714328afe4faec2e24df6e83fa2c38dac33f4302ebc89bddb4cd5f695321

Observation 4e9c196b-b341-4acd-9a35-ed535fcaf0b0 · outbound

This paper cites Axiomatic attribution for deep networks.

MIB: A Mechanistic Interpretability Benchmark Axiomatic attribution for deep networks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.095897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.421388Z digest=sha256:b2da918d0d19893f0601f21b050332b39165ccd6fcf1c722ab11e0fb3a24f562

Observation a48464b1-6400-44eb-81ac-bd3f8865dbf6 · outbound

This paper cites Attribution patching outperforms automated circuit discovery.

MIB: A Mechanistic Interpretability Benchmark Attribution patching outperforms automated circuit discovery

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.425522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.425522Z digest=sha256:b1d9a0d7fa3197ead35c33eafa8f3345075b216dc8c29d7468da4e37faafe73b

Observation 00e92e82-b680-43db-8b5b-562772f559fd · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

MIB: A Mechanistic Interpretability Benchmark Linear Representations of Sentiment in Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.431272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.431272Z digest=sha256:07249dc46b08b1ee44ab8b8617dd56bb25ebba03b1a9556eaf87598a006e3643

Observation 58d770e9-5288-47a9-a52e-95f406d24dd5 · outbound

This paper cites Investigating gender bias in language models using causal mediation analysis.

MIB: A Mechanistic Interpretability Benchmark Investigating gender bias in language models using causal mediation analysis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.435562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.435562Z digest=sha256:ac547327b08fa1d989b5a799dde851cc244bcae176040f519c22f7efe8fc1a32

Observation 4622cb2e-3eb5-48f1-97fc-3ea60751501d · outbound

This paper cites R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J.

MIB: A Mechanistic Interpretability Benchmark R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.072195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.440347Z digest=sha256:5502f5afcc96d53bdd162bf821016242bc4b8ff50647a74c729feb1954419028

Observation d145f84b-072e-461d-9725-16ea5ef63ab7 · outbound

This paper cites Answer, assemble, ace: Understanding how LM s answer multiple choice questions.

MIB: A Mechanistic Interpretability Benchmark Answer, assemble, ace: Understanding how LM s answer multiple choice questions

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.051671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.445962Z digest=sha256:09b7cb2700e857011f4288d837552ba43555865dc1e087be898b424ad0ac1ac8

Observation 4a6add0d-85b1-4107-8ff1-40ae29e543c3 · outbound

This paper cites an unresolved cited work.

MIB: A Mechanistic Interpretability Benchmark Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:19:45.033726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.451531Z digest=sha256:0a6d43b48bf065d656025fda61fadb74e4efc17cef877c6afb501e86ecd1227e

Observation a23db8f0-df8a-41c6-8786-2add4f54c01e · outbound

This paper cites D., and Potts, C.

MIB: A Mechanistic Interpretability Benchmark D., and Potts, C

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.021012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.457623Z digest=sha256:d04e047b8c43ca35b1ebdd5f6cab705f2c2dd6bb6979c5ba1efbfcd862f414d1

Observation e4670055-fc91-487f-9650-02c24615c8a9 · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

MIB: A Mechanistic Interpretability Benchmark AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.465708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.465708Z digest=sha256:9db2d698880e5fcffb535ca005b3f36111efa3516ead33eea44475bda1d912e2

Observation 810c448a-957f-4077-be25-c39aa297cc9e · outbound

This paper cites Qwen2.5 Technical Report.

MIB: A Mechanistic Interpretability Benchmark Qwen2.5 Technical Report

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.471525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.471525Z digest=sha256:2c93505544eed83a642eb77e66b836c3aa4daf974b61b3a1edd0cb20221b6e49

Observation 1c8f33e0-d25e-4d84-98eb-68b458c52098 · outbound

This paper cites and Nanda, N.

MIB: A Mechanistic Interpretability Benchmark and Nanda, N

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.476379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.476379Z digest=sha256:f01d8ea93a99468e97acb0d844ccb5cf4df139f3a5194dd1c71c10b401b7a379

Observation 0216a669-8cd6-4d19-80de-143332625da7 · outbound

This paper cites Interpreting and improving large language models in arithmetic calculation.

MIB: A Mechanistic Interpretability Benchmark Interpreting and improving large language models in arithmetic calculation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:45.000056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T12:19:44.480799Z digest=sha256:3cc27eb69f3ff227b438d9d7b2a3f7c0173ce6151690cb9de783645867299f97

Observation e8690118-a74a-46bd-b5af-241c19b7936e · outbound

This paper cites MQ u AKE : Assessing knowledge editing in language models via multi-hop questions.

MIB: A Mechanistic Interpretability Benchmark MQ u AKE : Assessing knowledge editing in language models via multi-hop questions

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.485555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.485555Z digest=sha256:dce5acf3733da0e92f01d00d800e6e64b94a48798f48382a91efbaaea0acf287

Pith citing papers

Observation 6366d1e9-2358-4b4d-a9ef-ed0ae288c6f6 · inbound

Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation cites this paper.

Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation MIB: A Mechanistic Interpretability Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:30.524426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:30.524426Z digest=sha256:d10b36bcebbda3c62f3a0e2a00fccf34e26ca3ec9dc35079f6e837232d335bbf

Observation 631dce06-3a98-4fb6-90c6-92421f0be3ce · inbound

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender cites this paper.

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender MIB: A Mechanistic Interpretability Benchmark

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:16.659440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T04:50:43.547037Z digest=sha256:1e759a3d640cc00288da4f70bea759b594765da56be2289e2402a41d4cea9963

Observation 0380bb66-0e57-4391-a54d-9191c5e8e551 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State MIB: A Mechanistic Interpretability Benchmark

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:59:28.649944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T20:53:40.666929Z digest=sha256:09aceda4f49ab9c9676079111f4f2922c5b856c0919ab50818d44575b80619bd

Observation 54662dba-3de8-42c9-b2ae-4e3f66a3b62e · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State MIB: A Mechanistic Interpretability Benchmark

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:59:45.181193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T04:59:11.877068Z digest=sha256:1df308dbefc5081a25899fa0d4effc6548f32ad73d5827ac1165a3474ca2a0b2

Observation e5a8f98e-83c7-4617-9496-6f0468b548d8 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State MIB: A Mechanistic Interpretability Benchmark

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:53:47.240269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T21:49:47.934339Z digest=sha256:a6cce16faf2e4efdd6526843412043f0b95eb4f49cf3b05ac272ab87b5aa1cd9

Observation a4ace506-8545-4824-b27f-8caf6bb7ed2f · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State MIB: A Mechanistic Interpretability Benchmark

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:50.228598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T07:46:41.159688Z digest=sha256:c2657c6c6f51ff707e11fa0808bc4fc5afbf02bef44a6db67e1557e141d80c08

Observation 47aad02f-4569-4385-a330-036cf92d48ec · inbound

From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach cites this paper.

From Circuit Evidence to Mechanistic Theory: An Inductive Logic Approach MIB: A Mechanistic Interpretability Benchmark

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:03:59.212575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T06:01:36.610373Z digest=sha256:c81a32bcf8cbc67cadfebea5e51c3a3cd1f8499191c702620df94a192852833f

Observation eb5dc46d-99ca-4240-893c-f772b8b110c3 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model MIB: A Mechanistic Interpretability Benchmark

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:05:47.045442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:603db6ce6e1244e964ad41ebff65157714fc110029002bba80848ba9d4ada088

Observation c16647de-c97d-4f05-95fc-d49785494fdc · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model MIB: A Mechanistic Interpretability Benchmark

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:35156389f2e7aa0d0b89593e669c748d6b5fa74d82d2bbc748a308ac530680ed

Observation 4222ef77-f29b-4ed0-a020-dd1b420116c2 · inbound

ToxiREX: A Dataset on Toxic REasoning in ConteXt cites this paper.

ToxiREX: A Dataset on Toxic REasoning in ConteXt MIB: A Mechanistic Interpretability Benchmark

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T04:43:07.056875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T04:33:18.794505Z digest=sha256:fe4523ba36b5e46e5d1f1f52486d3fe07a29688fde6de00c5a5f120a7b78459f

Observation b7fbbf63-7d51-48b0-927c-cab0efac45f3 · inbound

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits cites this paper.

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits MIB: A Mechanistic Interpretability Benchmark

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:43.457359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-03T17:31:22.683561Z digest=sha256:464adf6a698de55b6a8deac22fc14b2f81ef84f394c544af5ea30238bb17ee5a