Pith. sign in

Paper Citation Record · LEDGER

Transformers are Multi-State RNNs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2401.06104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06104 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:46:11.411181Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb1af497-7763-481c-b9cc-695444dde576 · inbound

Gated Linear Attention Transformers with Hardware-Efficient Training cites this paper.

Gated Linear Attention Transformers with Hardware-Efficient Training Transformers are Multi-State RNNs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:15:14.156799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T01:15:13.991219Z digest=sha256:e172b145d786746646516a57148b4321cbece9d052a6cbe4074e8fda977d9f62

Observation 7b21014a-8b4f-447a-a90c-9c9e34f39312 · inbound

Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference cites this paper.

Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference Transformers are Multi-State RNNs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:12:20.906064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T14:12:20.809594Z digest=sha256:f7286fb1e4e05dbf8932381d8a1c1f9d4b9af4ff2600c5c3af19b3dcef4a380c

Observation 92e6b6c2-8183-4301-8eab-e3804da67195 · inbound

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU cites this paper.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Transformers are Multi-State RNNs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.902581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.902581Z digest=sha256:6b9c429ffd6cadca801ddc7e7e166969bddd21b8ba491365be8a1c8b8832e19a

Observation 00232cfb-b8d0-4ffd-92b7-ac3239210597 · inbound

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification cites this paper.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Transformers are Multi-State RNNs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.411181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.411181Z digest=sha256:312dd276eb9486db6ca772c95ac832c83c2068f3da12c0ff82d08737be9936ad

Observation f8187306-dc17-460f-b73e-463d1aec41f5 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Transformers are Multi-State RNNs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.269719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:fc092dd2c298a369fd964f1e1a6c35232ee4a43a96eb4324814758cf5f8094e6

Observation c33b358f-77c5-4223-bb69-6eaff1831127 · inbound

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression cites this paper.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transformers are Multi-State RNNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.009571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.009571Z digest=sha256:976a460a9505a63d42aa2528e8ec6101e0f112c84f8cb321f13816d7fbac0f36

Observation acd58307-ba1d-475d-989c-cde40fe46f23 · inbound

Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models cites this paper.

Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models Transformers are Multi-State RNNs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:42.533672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:42.533672Z digest=sha256:eae563068ee71fdf754ccbfc181a7af31aa59b15e6817d9a938bf82c0728ed4e

Observation c0e8cf90-f4e9-4d22-b575-2d56dd548708 · inbound

AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models cites this paper.

AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models Transformers are Multi-State RNNs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:05.812475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:05.812475Z digest=sha256:fd5930f641bc1c5e6d305d398c29ac7c2b46ddb18b6cc230d4f0f929bac0ddf5

Observation c4e08f49-c96b-4585-bffa-050c6b1f3ab3 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Transformers are Multi-State RNNs

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:33.321894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:33.321894Z digest=sha256:4f95d08edfed980b55bc338282ed37a3f613716358dbeea183075c4773689844

Observation 68a3ec76-a728-43d3-a209-688f3381f75f · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Transformers are Multi-State RNNs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.564246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.564246Z digest=sha256:1a6dd21d5b8bb0acf2b4331d686d538ee8143327af28a72d7d25af121b75651a

Observation 99117483-cd98-46ce-bc6c-8348b4157fd5 · inbound

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models cites this paper.

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models Transformers are Multi-State RNNs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:33:20.696069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:33:20.696069Z digest=sha256:ad0a5622644d73a3ff2f37356633d95ddb549cdb412aa24faf228f1eae04d618

Observation 9a51bea7-5360-48d8-995f-7b83e8aec3b8 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Transformers are Multi-State RNNs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.525852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.525852Z digest=sha256:8b6d8acd06ebcd7d0d175ea220d65461f38516e80d53eda0f4b20f66c6ec3e63

Observation bc49c771-0dae-4a86-a69e-05b535178225 · inbound

The Pitfalls of KV Cache Compression cites this paper.

The Pitfalls of KV Cache Compression Transformers are Multi-State RNNs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:32:35.105944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:31:57.865691Z digest=sha256:33c1aa09890b0ed84448898222d15bea63f13bda41aeffee4244d7c5500c82bf

Observation 0da759d6-e3a3-47f4-8bed-dee80839d12a · inbound

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference cites this paper.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Transformers are Multi-State RNNs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.739935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.739935Z digest=sha256:fcc51c5c62aa59ec961cc0bd713c4b94b33c49e0a6cc75ad2709e08079d6da6c

Observation bbbe9cf3-856f-4f15-ae92-43e88bceab4c · inbound

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning cites this paper.

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning Transformers are Multi-State RNNs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:26:02.911734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T07:25:56.953876Z digest=sha256:c25236dbffd7e61a9424af7b734c586f68b8dd608ff73cac39be9c383fb94366

Observation 5ebeb06c-dd57-482e-9dd1-c700f1f3d813 · inbound

Learning to Evict from Key-Value Cache cites this paper.

Learning to Evict from Key-Value Cache Transformers are Multi-State RNNs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T01:18:04.039897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:18:04.039897Z digest=sha256:1b34e83cb7aacd9716cc8c7c27a6db2de485ccbce8d2fbc872ac87405969d3cf

Observation 493ee67c-1b0b-4b32-bb7b-885bdd991d67 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Transformers are Multi-State RNNs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:00:17.921407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:59:33.902420Z digest=sha256:0f979002a3789cdd70f2b22bdcbf7fe34fdba03fc898a3b9931796e2d0cc41b5

Observation ae59597a-eb62-4a02-b206-5c0b9a19dbd7 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Transformers are Multi-State RNNs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:50:09.524907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T12:45:27.150368Z digest=sha256:a3aab72f9c4872301b4ccee0045a8b73b21ed011bca8f62aef771b15a23a2b6c

Observation 1a5cab8b-0878-4fba-aa6d-38b027148999 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Transformers are Multi-State RNNs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:05:43.031477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:05:43.031477Z digest=sha256:3d71e75f15a1e9c1e93fc7a11fc2fab759675ce959095f4f9e9fa577f2ce951d

Observation ccc1e841-1654-4da3-bcbf-63cb8c3123b6 · inbound

Stem: Rethinking Causal Information Flow in Sparse Attention cites this paper.

Stem: Rethinking Causal Information Flow in Sparse Attention Transformers are Multi-State RNNs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T02:39:29.428722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:39:29.428722Z digest=sha256:add1df5af3d2371b41fce5ba9af190c19629744d57270e9e4581ebb45233ec6c

Observation 74978e6a-de28-433d-a557-e81d6837965f · inbound

Transactional Attention: Semantic Sponsorship for KV-Cache Retention cites this paper.

Transactional Attention: Semantic Sponsorship for KV-Cache Retention Transformers are Multi-State RNNs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:10:30.674922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T15:10:24.725639Z digest=sha256:f949e9ac25a6126143ac82db873a3a91f4772726e79db48fdad65f6d13badc1e

Observation 68e5b953-cc61-4190-9508-23742921cc3e · inbound

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization cites this paper.

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization Transformers are Multi-State RNNs

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:06:04.533339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:09:11.684432Z digest=sha256:2f6b6dbb27915706a0b979397e16ba643ad25bceb3d83c30b004cbc2c9e74ebe

Observation 7d39eab4-515e-473e-84d7-78f822fabc8f · inbound

Long Context Pre-Training with Lighthouse Attention cites this paper.

Long Context Pre-Training with Lighthouse Attention Transformers are Multi-State RNNs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:09.305982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:10:38.610613Z digest=sha256:d900dccca52e4cf4c91a7b506f40569f2757b7296b09cc6e0cea3a2d40d1d3b7

Observation 3c6a7848-cc99-4c76-a756-5e0081d6b7d6 · inbound

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing cites this paper.

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing Transformers are Multi-State RNNs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:26:30.124036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:52:33.076123Z digest=sha256:e6f0a08f42e157019e15f274533dcf747ba82ebdde1394f12c5c1df195aa9311

Observation adbe79b1-3ddb-4618-b1fa-2d23c66e2ec8 · inbound

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference cites this paper.

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference Transformers are Multi-State RNNs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:03:59.894407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:03:07.685253Z digest=sha256:5c1ae2dd3732fb81c579bd87ab902120c37f4759c96621516e24ecb171ea7f9f

Observation 716d21de-9384-4131-aa96-0b597c708455 · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling Transformers are Multi-State RNNs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.380935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:76765e92baf5521adf76cbba8748a68f185b30a2ab5594e5c10d584c1af7a112

Observation 075db97b-afd7-43b0-ba76-71eb3d78ff6e · inbound

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding cites this paper.

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding Transformers are Multi-State RNNs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.011184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:11:58.636939Z digest=sha256:3430c25d6641f9c72455d3fa1bfb534d0d112d0162ad36d00decc31ca640c976

Observation e953ea22-6f25-489a-86e2-59e1f78e8454 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Transformers are Multi-State RNNs

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.005016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:e5c93c17afccea2a569fefc5308ac1f14ce0b0433355f4967e1299e0dc921c58

Observation ea27f054-3853-440c-a141-84780c5d90fd · inbound

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference cites this paper.

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Transformers are Multi-State RNNs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T10:41:08.967261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:41:08.967261Z digest=sha256:1763605e34c53bab48f829206ec18d03ed721f8c7ebf4aab63fe7ba0351e0688

Observation 5b2fe512-ee50-4301-acee-4e5e2a4cda4e · inbound

Error Certificates for KV-Cache Eviction via Randomized Design cites this paper.

Error Certificates for KV-Cache Eviction via Randomized Design Transformers are Multi-State RNNs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:23:39.162227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:23:39.162227Z digest=sha256:29f6c59c90db05d5d01f33867bd0649cb44aff133e192dc32f5cd6a8349e2d1a

Observation 21c45e05-256e-4f45-aed1-31751a9dd61c · inbound

Training-Free Hashing-Based Attention via Binary Principal Components cites this paper.

Training-Free Hashing-Based Attention via Binary Principal Components Transformers are Multi-State RNNs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:48.501586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:45:48.501586Z digest=sha256:e410980946613c8ae1d63564c92b1b4213498fa88786b4aaa7a81174f604a3e8