Pith. sign in

Paper Citation Record · LEDGER

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 6 inbound Pith citation observations for arXiv:2504.20938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20938 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:21:34.095808Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:40:01.662137Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 841b2460-9333-4104-bae8-db0db3b94347 · outbound

This paper cites An X-Ray Is Worth 15 Features: Sparse Autoencoders for Interpretable Radiology Report Generation.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition An X-Ray Is Worth 15 Features: Sparse Autoencoders for Interpretable Radiology Report Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.854161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.854161Z digest=sha256:8c04121822c9065718db1ec37be4973bdadfd9ef5056f7547406f28c9e447b57

Observation ceae6dea-c0d5-4251-a414-04ee162b4c3d · outbound

This paper cites GQA: training generalized multi-query transformer models from multi-head checkpoints.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition GQA: training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.859560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.859560Z digest=sha256:4d77db85d0c83e9476a02b2f4c82afbfddc323380fcf9f520ec4ba3eb5c7b6ec

Observation df625bae-2760-4f66-8a81-9a207731c4b6 · outbound

This paper cites Investigating successor heads.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Investigating successor heads

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.769922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.864209Z digest=sha256:2445df1011e0024245b7dcd292be3e5ec8d5984661babfcb78d06dea90aad7f6

Observation 9b3906e0-73a7-4a95-a1a5-bca9a937d9ed · outbound

This paper cites an unresolved cited work.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.869155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.869155Z digest=sha256:4ed9331cfa2cb9ea28f5cdb0f1f85c9cad0dbf425cad1362362d2af1505b602d

Observation 203a3c23-04fc-456c-b8b8-4930711d7017 · outbound

This paper cites Linear algebraic structure of word senses, with applications to polysemy.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Linear algebraic structure of word senses, with applications to polysemy

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-08-16T05:21:33.873953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.873953Z digest=sha256:9f1286815b69540c127dfb12bdace07281aa07237658b6b98a18f38d1f171505

Observation 81df82d4-9c4e-4d37-a38f-835d1e0ca6b7 · outbound

This paper cites Circuits updates - march 2024.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Circuits updates - march 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.747690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.878910Z digest=sha256:d172b920f32bfcad46913d9e1c30a6ef7d662055d339433b8454b56465f158ae

Observation 54b5c374-4026-4ff3-ba22-37e2805962d9 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Pythia: A suite for analyzing large language models across training and scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.883448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.883448Z digest=sha256:4854ea613083bf5e9949759462d8d7b406e20cd702c4b0b42f5fa1ef83556e34

Observation cea670da-0da0-40e5-9a47-e225d88ec91e · outbound

This paper cites Language models can explain neurons in language models.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Language models can explain neurons in language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.888631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.888631Z digest=sha256:64872c385d93040b664c45faa2f9098926aad1a9ffebcadb34f03e371cbec4e8

Observation c046db1a-c4f2-429a-b961-8591ed0d48e0 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Towards monosemanticity: Decomposing language models with dictionary learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.893071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.893071Z digest=sha256:fbe3094b74e3feb1f286f928ef8a30cdf46f23be6e0d51a9aac1e611130cdd04

Observation 3e5a0caf-c503-46dc-84cb-59fb8b7be571 · outbound

This paper cites Learning multi-level features with matryoshka saes.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Learning multi-level features with matryoshka saes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.706877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.897440Z digest=sha256:3213238b939b77d7fa3c70648bc9e9a47b7f123326c8a188f6fc27429e061650

Observation 1dc1d14a-2357-4533-93e5-d22d75f7fd20 · outbound

This paper cites Circuits updates - february 2024.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Circuits updates - february 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.692899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.902499Z digest=sha256:8293352ad54af73b0de2bee53f3b818597e1c053f8167a8152fb9adf1538a6c9

Observation 6eb2bb85-6b4b-4f0a-b88b-affe2090c58d · outbound

This paper cites Circuits updates - april 2024.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Circuits updates - april 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.678910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.907898Z digest=sha256:c98348bebfe2c36bf57b25101f2d51b4ceba4df30bed42666d7ee90ff9d473c1

Observation ae1b040f-4853-49ed-819a-e31cf022f2b3 · outbound

This paper cites Mavor - Parker, Aengus Lynch, Stefan Heimersheim, and Adri \` a Garriga - Alonso.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Mavor - Parker, Aengus Lynch, Stefan Heimersheim, and Adri \` a Garriga - Alonso

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.664772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.912701Z digest=sha256:6c2449a16a6c15417807e8cdb48e939d0f472ccc03dc6ed4e5bb26ffbdcfa837

Observation 91d223b7-b1ee-4258-99fd-e118d10e9ef1 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.917333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.917333Z digest=sha256:f14d0825e6d7cbd5429e6a5e96125be9631e648cd923b7fb35f0394f38e66cbf

Observation ec24089e-a856-4782-9c6c-fce04bad0948 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.922838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.922838Z digest=sha256:55f9f7db00afd427efe27de125fa625d3689297e985f2970a2438b77d8e8e597

Observation a5bf3096-13ae-4fc6-b63d-faf210e7058f · outbound

This paper cites Transcoders Find Interpretable LLM Feature Circuits.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Transcoders Find Interpretable LLM Feature Circuits

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.927686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.927686Z digest=sha256:1390f852852de5bc2fcba703fa5b8a61761818aa025ca3d052bf6288980b3ec6

Observation fc86a35e-f928-417a-a7c5-5ec4e77a78ca · outbound

This paper cites A mathematical framework for transformer circuits.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition A mathematical framework for transformer circuits

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.932861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.932861Z digest=sha256:498d579aea7ad1b0e3f87eeb007cc5b87f7304ca4d83a62af3066206aa498f0f

Observation ebaf0d76-001b-4397-bda0-f42e89141370 · outbound

This paper cites Toy models of superposition.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Toy models of superposition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.641060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.937511Z digest=sha256:1563f64566d52f53f1c6494d75435c2e00135998027375967931d6c9e004ace3

Observation e16a0c9f-c04d-4537-902d-0ce1ef1cc6e7 · outbound

This paper cites Privileged bases in the transformer residual stream.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Privileged bases in the transformer residual stream

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.627598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.942008Z digest=sha256:8f287177c10ca901b776eac30ccc594ba303ea32a4815e5437402351265f2be2

Observation d9265d54-85bc-4c3f-a5d8-568997fd9610 · outbound

This paper cites Decomposing The Dark Matter of Sparse Autoencoders.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Decomposing The Dark Matter of Sparse Autoencoders

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.946749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.946749Z digest=sha256:6b64dda4db9b205ccd3c57f67dec865b2327a3a44d0f10d608777b3070f3cdf4

Observation ff1f572f-93e7-43a3-9510-45cec4ace6bd · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Scaling and evaluating sparse autoencoders

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.951798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.951798Z digest=sha256:e8c708d15365c1fb50d653f4d01e6d309e72b0fe966030f68477e04220af5fec

Observation d6cf6feb-bd66-40f0-a75d-681ee4e6d83e · outbound

This paper cites Automatically Identifying Local and Global Circuits with Linear Computation Graphs.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Automatically Identifying Local and Global Circuits with Linear Computation Graphs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.956907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.956907Z digest=sha256:7811c7536289efab8028a190b7f469edbf3832ef237c22ac515a0c44979c780f

Observation a58857c1-5ada-4196-8028-fde66498b802 · outbound

This paper cites Successor heads: Recurring, interpretable attention heads in the wild.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Successor heads: Recurring, interpretable attention heads in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.612730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.961753Z digest=sha256:0713b2a7149def509f976565eeed3603403250970a3b948beadb5eab4d8ae9bf

Observation 0ab48272-ce68-4d3b-a8b8-b70b03f3d1db · outbound

This paper cites Finding neurons in a haystack: Case studies with sparse probing.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Finding neurons in a haystack: Case studies with sparse probing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.596174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.966059Z digest=sha256:2ca01e9eb6df2bcb521503f2776976f4604d5059132b5bb5a7eccf2e10fc9211

Observation bd92c09d-e49c-4120-a465-a4e346d61189 · outbound

This paper cites How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.582097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.970624Z digest=sha256:87f917537353d82c8b7f47ca79916b29acd708880d271a1b5f30dbde1b581372

Observation 818b2bf1-2217-4704-8910-06b8fca051c0 · outbound

This paper cites Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.974837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.974837Z digest=sha256:aeb3274e15bf0cc12bdad18c5d4ec54d540bb6e9fbca10acf39b30e2b445d5f7

Observation 1859745b-79bc-462e-a3b7-664f1a81a826 · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.979344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.979344Z digest=sha256:a0369e46d7370fd0aa70aa68e0b501528b3c554f5b811594b70771bc21bbb7bd

Observation ba8cdf98-03f3-4ef8-97ed-0e872d23090f · outbound

This paper cites Circuits updates - january 2024.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Circuits updates - january 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.568332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.984871Z digest=sha256:824f7d05a047e46de8484d2e88de6116c97283c51ee10f429c8503a6cc0061ce

Observation 7f224a36-a82f-44e2-9f41-d3d1094094d8 · outbound

This paper cites Interpreting Attention Layer Outputs with Sparse Autoencoders.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Interpreting Attention Layer Outputs with Sparse Autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:33.990706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:33.990706Z digest=sha256:03309d0e3fb374e016de0fff0e305935009b4f51470a2c21aad4fee3ec4b752a

Observation f3613b24-d840-4eb9-aaf2-c378690cef2d · outbound

This paper cites We inspected every head in gpt-2 small using saes so you don’t have to.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition We inspected every head in gpt-2 small using saes so you don’t have to

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.554227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:33.996065Z digest=sha256:5cdbf398f68bb9c77304e9f22bf3268c1101045deccc22ba24e1ab96230ca7e7

Observation 577fe07b-c46c-48ea-a8ea-f9494129aaaf · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.000670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.000670Z digest=sha256:e2f006cbedeb29ee59346e36579ecca9ffe7d294e2b24b2c0b4cfdd0354d616c

Observation 1aa8995e-6ccb-4ff6-b9e7-ab4964fc8759 · outbound

This paper cites Sparse crosscoders for cross-layer features and model diffing.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Sparse crosscoders for cross-layer features and model diffing

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.539644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.005437Z digest=sha256:e4dd98fda4ff310d3836151b43d8cb0014b648b3a31c377c4ae1ad5727136289

Observation a8c42e8e-1bde-46a3-b6c7-2b9908ec22f2 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.010473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.010473Z digest=sha256:d7791c27c14084d5e383af9bbb08668c93ce672c0dfe4c3a325d8710cdbd87d7

Observation b85343e4-9dc7-4f1d-8b7e-606bd1ed413d · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Copy Suppression: Comprehensively Understanding an Attention Head

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.015126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.015126Z digest=sha256:f2674817753aa5e96c9e2a3b7c477adc1fe974b5f6b79b67a01ee1cdfebb7fac

Observation 97b2d602-dc0f-4d0b-995e-5da403c2a3b4 · outbound

This paper cites Locating and editing factual associations in GPT.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Locating and editing factual associations in GPT

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.019977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.019977Z digest=sha256:1f112f888601001568eee7837cace39de019123d1b3d03757029d5068727d8c6

Observation d997c440-e44b-4baf-a679-a3adc47c2fce · outbound

This paper cites Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.024650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.024650Z digest=sha256:a8ca8552a32466a98ebe490377140197034840bad0e33efccbb73ddd1c0c43c1

Observation af6c29e5-309f-40ac-9a61-3cb047bf91b2 · outbound

This paper cites interpreting gpt: the logit lens.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition interpreting gpt: the logit lens

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.507329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.029209Z digest=sha256:f15cef82fe19dcbf6a8543d920acb15aeb29c651ff69afa7775878528dc0c473

Observation 98d853a8-1720-4094-842b-dc2fe2656133 · outbound

This paper cites Circuits updates - july 2024.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Circuits updates - july 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.492898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.034720Z digest=sha256:166ab82b45f7a087534ebc1ffe8a07afba8812624521294c6e469c0289ef843d

Observation 93f757ae-d6f1-4fc3-9727-99f21e847a42 · outbound

This paper cites Zoom in: An introduction to circuits.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Zoom in: An introduction to circuits

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.038983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.038983Z digest=sha256:44bb779b188bf52dfa46588a5e700e834286baf9d74c3e83df265f0ced43e085

Observation 00284d91-c70e-4c63-bf78-a999c48d104f · outbound

This paper cites In-context learning and induction heads.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition In-context learning and induction heads

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.478775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.043226Z digest=sha256:955818efc75e4f00aeb1b7bcdf66afc45ef3d7a9101143da969081acb99d7870

Observation b248592b-186f-4cab-9fb0-9aa1d116428f · outbound

This paper cites Improving Dictionary Learning with Gated Sparse Autoencoders.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Improving Dictionary Learning with Gated Sparse Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.047569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.047569Z digest=sha256:d7609b3362a4635c8a77211b4f412fb86148fad7bd117aac7e9aee71ac81a8f1

Observation 79db3ea2-c12f-4cd1-a85b-fd8073caaf59 · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.052412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.052412Z digest=sha256:e165192ca5121bfcc580015fa362f3c8d41da5a0d293e8e9a304e297e3f82bb1

Observation 271d5937-0142-44a2-8d91-b17868e71768 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.056864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.056864Z digest=sha256:157a5c4c110cf7f4a5793d941b29bab1128a6c839c5e4a50421e85faa8006070

Observation 9b1f6174-bb1e-4d50-9305-1de5845481cc · outbound

This paper cites Circuits updates - january 2024.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Circuits updates - january 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.454542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.064399Z digest=sha256:a45f196dc3f8a84900a679098a8f5abcb7328c041213954e20c40b54e75f1ceb

Observation 1e307001-eaf0-49bd-a53e-be0e8e2240ff · outbound

This paper cites Daniel Freeman, Theodore R.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Daniel Freeman, Theodore R

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.440784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.069289Z digest=sha256:9e3b546ca95493a4cac0ff32bbb4747b2f2ed03801ffcb9a0562c8f1615682bc

Observation 529f19ca-c0c0-4510-ab3e-8d572bbe242d · outbound

This paper cites Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.073745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.073745Z digest=sha256:ef29d2870e811815abc46f933aa04c111543fd7bcae028bdaad4ca1dfb2958da

Observation 0c8aa047-1c47-4bb0-9e99-39589cd655a7 · outbound

This paper cites Interpretability in the wild: a circuit for indirect object identification in GPT-2 small.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Interpretability in the wild: a circuit for indirect object identification in GPT-2 small

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.078512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.078512Z digest=sha256:c5e3502a5f900b8cf6e2ec7feb9aa10c7873c57f759c0ed834ac5dd6f2d02b11

Observation 24d7cee7-9a87-435d-b255-491aeacf4fb5 · outbound

This paper cites Addressing feature suppression in saes.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Addressing feature suppression in saes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.416558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.082779Z digest=sha256:33f6d02151c7e9ad8c68dd1dd1f0c46b772de1157a630bd225338a70f45c828d

Observation 68f4bc23-0f97-4463-b84a-2221f2d7af7c · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.087163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.087163Z digest=sha256:84532bc9e6893ba49655ca7280f59c05da2b80ede71ac8b597d05662581755bd

Observation 209b7b98-1cc0-43a8-b31c-47ba5881e880 · outbound

This paper cites Efficient streaming language models with attention sinks.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Efficient streaming language models with attention sinks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:21:34.091712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:21:34.091712Z digest=sha256:466c0742ab36d1f02a8fa9c1295b14984ac6b2e47b0797c3c8f11ba85b101486

Observation 3b9c9586-66af-459f-aa6d-ef156412683a · outbound

This paper cites Towards best practices of activation patching in language models: Metrics and methods.

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition Towards best practices of activation patching in language models: Metrics and methods

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:21:34.392843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T05:21:34.095808Z digest=sha256:a5700b7b9d915c33ea2e0c2d0673ae42a313719b0dd56467bb0e67637b304581

Pith citing papers

Observation c39b9499-16e6-4950-b3df-ba9b0e5c3b20 · inbound

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping cites this paper.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.846244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:878d7cf9a7af46d263a82643127d34c55d79cbc9384055dae71d520d5c7bcdf1

Observation b5ce9711-390c-4319-9681-f6ab24dbc286 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:28.658104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T20:53:40.666929Z digest=sha256:dff550af7ddde43d98fecada9dc3efd9c85807d1c5205afffc8fa7a148f1c054

Observation 4d84dd86-ebbf-4ee0-b0c9-3a182183b521 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:45.258995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T04:59:11.877068Z digest=sha256:bf3e7c6f92324add12266347f0f673cd562c77e67e691682d23e4d2e989d0e75

Observation 328a9f07-b921-4320-8a41-b73b6a33b28c · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.682964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:9785800db74e53fb78ace6292ca9036e8c38d5fc6c629182ca9ffadcce52754a

Observation d0531adb-c723-4a4c-aa41-e2b524fd2c5b · inbound

Targeted Recovery of Weight-Space Mechanisms From Neural Networks cites this paper.

Targeted Recovery of Weight-Space Mechanisms From Neural Networks Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T10:39:48.151675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:39:48.151675Z digest=sha256:c65a8d365e4f8e247d8b9a4eda75815c7057f1bb685585495df53008deb9be31

Observation e9b28670-8a8f-4615-a2f0-4485e8e3997d · inbound

Targeted Recovery of Weight-Space Mechanisms From Neural Networks cites this paper.

Targeted Recovery of Weight-Space Mechanisms From Neural Networks Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-02T10:40:01.662137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:40:01.662137Z digest=sha256:88cd4d5b3ae637de6c4a81aa25d494d47c35c697c80608588d38e0db66d5e74a