Pith. sign in

Paper Citation Record · LEDGER

Repeat After Me: Transformers are Better than State Space Models at Copying

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2402.01032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01032 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:05:31.626283Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:39.775848Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4785e0e2-ffce-4e15-87fe-1fa1c43050d1 · inbound

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models cites this paper.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.519265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:dad1fe241bd6f17a5302a571cabab63a3a39c4e60c557680bf15dfb2434ade0d

Observation f6b2ad84-3d20-4dd8-bc39-137bce7e2a5b · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:31:03.861393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:639bb97b9de3a7bcb4b39f2f1e1d786d7a135681550e56bedae9e0ceb5d61669

Observation 7dfd5ca6-f645-4169-bae3-35077d083a7e · inbound

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach cites this paper.

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:05:31.626283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:05:31.626283Z digest=sha256:1c4e0e2fb564465a01a5874274d97efc26a88198d99fc2720e2235cf818423f7

Observation cc935757-5111-4368-955a-85e43c483e88 · inbound

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid cites this paper.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.671185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.671185Z digest=sha256:6dc5e54fbbbe11cc8317c62c02de40d04150b26ab13a33f9d2986353798cddaa

Observation cc493380-73bd-476e-97c9-16df01d2ce48 · inbound

Sparsified State-Space Models are Efficient Highway Networks cites this paper.

Sparsified State-Space Models are Efficient Highway Networks Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:45.540484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:45.540484Z digest=sha256:943a88312275f34f6b7ea5d690a92ac44e28fe7d7217e720ab0bb7caa73f80e5

Observation 7d3ab2f4-6822-4c76-afe7-47e5270748c0 · inbound

Learning Compositional Functions with Transformers from Easy-to-Hard Data cites this paper.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:36.145807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:36.145807Z digest=sha256:c2f6da5b14d8e39107985965cccd70ca9e81edf73f34aaf0d4405b7085a9c744

Observation ac4be71b-5e71-480c-9044-2122935aa254 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:38.690738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:38.690738Z digest=sha256:2b5112267fe9aff59772bf8c5b312bb89aa64d22b0985291d9c7092a4e7afad9

Observation 9ae3ec0a-4294-471e-999e-c31d90f76e51 · inbound

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling cites this paper.

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:59.785532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:59.785532Z digest=sha256:a36a76f83acfb023c792c8042a92ac2d781e7c41700052b634b474ff84176b49

Observation 5e71ee86-8927-4fc7-b8d0-87a84a262889 · inbound

TTT3R: 3D Reconstruction as Test-Time Training cites this paper.

TTT3R: 3D Reconstruction as Test-Time Training Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:41:16.436769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T06:41:16.306593Z digest=sha256:49158b687ccdf01dbff6b497d9481f01e3e1231c245bdf3f1a3e7be592fb6be1

Observation 4ec361f8-de23-415e-8b48-2abadecf5c13 · inbound

To model human linguistic prediction, make LLMs less superhuman cites this paper.

To model human linguistic prediction, make LLMs less superhuman Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:17.799547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:55:17.799547Z digest=sha256:74a811cb61a94ca2aace7de535cc8449a984e40955f437532ebd79404c209fb4

Observation f3f7ac2c-eccf-44b4-984b-62f8888dbd4a · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.163140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:22cd5ea9767364489168d20ce085a6f2ba8b4efa71077b1d246bb92ddd9fac3b

Observation 8103b078-7df7-4329-8404-fb9a9755987a · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:56.548151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:56.548151Z digest=sha256:659236f1261aebfc25c025dff3c80dcebec916c2f38ae809f6147fccb16beeee

Observation 7cd71372-04b3-411a-86a3-38862c2577d6 · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.090936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:0f9af06fff875944d83632c6088e070a9b9d1dd598985e7dd067f48d34076a1e

Observation 31bba847-f7ed-4acd-9271-7a3a672b259d · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:50.161253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:50.161253Z digest=sha256:d8a49cbbde43e6c0e894df731b0e6814bc2e6229577d7ba955e0a6c9eff79e3d

Observation 31684462-8ffd-4905-adcd-93b9b894d592 · inbound

The Bayesian Geometry of Transformer Attention cites this paper.

The Bayesian Geometry of Transformer Attention Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T16:20:20.659781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T16:15:57.408685Z digest=sha256:810d573d7bdaa8997d299c68c082e22f9695f35be74234142edbba83d0e6730e

Observation 52bab556-5e0a-451b-974f-9fde3c04def9 · inbound

Towards Understanding What State Space Models Learn About Code cites this paper.

Towards Understanding What State Space Models Learn About Code Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:51:43.479561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:51:43.479561Z digest=sha256:97b230198aacfd9bb957c406ec0c81dc9bc254b692f78af13f01658d9b0996fd

Observation effe2026-80e9-4099-9641-bacc9a687709 · inbound

The UNDO Flip-Flop: A Controlled Probe for Reversible Semantic State Management in State Space Model cites this paper.

The UNDO Flip-Flop: A Controlled Probe for Reversible Semantic State Management in State Space Model Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.709421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:45:35.405345Z digest=sha256:bc5d536177726e0247b8d7b6aa628527e09f782dd2fc2d6a32abb6f757049f40

Observation 5215fe50-c4af-46e9-bf7e-663a5d4856ef · inbound

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding cites this paper.

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.627239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T22:03:35.260373Z digest=sha256:4a09a4a81d8442e12f561c5d9d01c5ed6e3c9ebd7b950da99bb2ff9ee25e321b

Observation a966edb1-f8a6-4fb8-b3dd-aac28ccb0c79 · inbound

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models cites this paper.

Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:56.203904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:07:33.756985Z digest=sha256:0af213e25f6130fff0d80eca0d4512a97004694dcf898ad4b5b53e7ac3d87c4f

Observation ad50ff5c-3d7c-4fe5-bd79-3d8a57f0aca5 · inbound

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators cites this paper.

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.325867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:26:48.026538Z digest=sha256:644bbccb574bc3aaa9c979b6f80d975a5b96b3910cc44ad751be113e6c4093d9

Observation bc1e3a80-bddc-4fb0-b861-9bad26bb08db · inbound

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases cites this paper.

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:27.007244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T04:29:37.989500Z digest=sha256:0fbd5ea62926a8167efae3379694a28953caef438d0a673070d87081d301257f

Observation e827a593-b678-41e4-b8a5-15afeea5ad62 · inbound

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention cites this paper.

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:09:23.228532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:08:18.768344Z digest=sha256:0c7bac604f841bbc1ff15bc22a1f9879de19eb7b7a194c1b4b62ff4fd933c6c2

Observation 7606210a-7535-4e13-a8d3-435571d897c7 · inbound

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference cites this paper.

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:43:59.528847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T21:37:51.638904Z digest=sha256:79c83fada8f50d707c67b72c6ad620ef30bf0f7e8c8639384f418df29baa0ed6

Observation f65382ec-1ece-4ce6-9e6f-5f4802358b7d · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.673759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:de667a2e5f48086da744c2d7403c3975238ab905f9c4bfdb1d9226e6ada3ae8f

Observation 3c7b0e8c-ea93-4311-8fa3-f457973d9ff0 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.365982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:9b1c34d32ac345b09eeaf6a18e36ff7b2d802753adc616082367d66243845d02

Observation dff055e0-ab10-4f46-938d-37dd97af1876 · inbound

A Verifiable Search Is Not a Learnable Chain-of-Thought cites this paper.

A Verifiable Search Is Not a Learnable Chain-of-Thought Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:39.777090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:35:20.698118Z digest=sha256:4a7c5e12daa5158dc9505d5563762ad36576443d72b48aef6f4aa75c57ba09fd

Observation ed936055-c6b7-4404-84e4-ce80ef412ce5 · inbound

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets cites this paper.

A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:48:20.316429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T13:46:18.862925Z digest=sha256:afcae164d54dd89f16782c930e7074228a400ee86f4edf863ceccbacd7185007

Observation c975545d-840b-4086-ad3b-027c90be10fd · inbound

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention cites this paper.

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T14:50:03.831572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:50:03.831572Z digest=sha256:c6457782a220f59afc1e861e83befaec8f8e417e7f0b67be7cdfd16ad746d094

Observation 0aea3c71-0acc-4a0a-831f-6f5ac09bd726 · inbound

DSSMs: State Space Models with Explicit Memory via Delay Differential Equations cites this paper.

DSSMs: State Space Models with Explicit Memory via Delay Differential Equations Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T13:15:38.524121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:15:38.524121Z digest=sha256:d993f1476daf2a558b77dc035300319aef04a689912a5246dd07c07326489fbf

Observation 1ce7cf71-0aa1-429a-acce-aae0ba99215a · inbound

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams cites this paper.

Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T11:45:09.424370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:45:09.424370Z digest=sha256:e862b258f01c0a6c42b407c5e90407edabb6cd0c2355b1500a47110d32c59d0b

Observation e90e7497-c8d5-4d09-bdfa-ee34ab6c017d · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:39:22.140957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:39:22.140957Z digest=sha256:c7e43bc7fbfecff1ba06180592ed2a346c6e49539608074e13464f9e35da248b

Observation a3aa00d8-fc6d-4ebf-b856-35e685c3f74f · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T02:03:24.068474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:03:24.068474Z digest=sha256:b7b2665ec1ebe04a1217e7b6fbce5f7bdfd12e04e2097d3acab54526099598a9

Observation 86d98c18-52eb-4932-92da-eaadab7a7c71 · inbound

Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones cites this paper.

Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:06:48.827702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:06:48.827702Z digest=sha256:8a553106fb643f9b2ade9aae1e716448e67406e4f2c9e326a87949e266b16e6a

Observation d26c5665-9749-4773-b242-650aa8297dc6 · inbound

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks cites this paper.

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:31:47.771481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:31:47.771481Z digest=sha256:bba5157a4ea7117f38284ac18138f42a4fae323eecc4d32eaba01d320d70bdb9

Observation 653d6092-2031-4eaf-9b84-61e86a5c15f9 · inbound

Raven: High-Recall Sequence Modeling with Sparse Memory Routing cites this paper.

Raven: High-Recall Sequence Modeling with Sparse Memory Routing Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:01.258380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:01.258380Z digest=sha256:1a77b630c2a059f5ed1968a4b5ac65359e3b7f01728e1d3585dcfb639081e43d

Observation 74de6bad-a1d0-4c22-af69-25568eff1569 · inbound

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling cites this paper.

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:14.385860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:14.385860Z digest=sha256:941e3628931ae728baea6e25e866594dbd597b05f6d02dcc6d1c98a36b236cae