Pith. sign in

Paper Citation Record · LEDGER

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

As of 13 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2501.00658.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00658 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:52:33.772825Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:54.578841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T04:25:56.146025Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db420ba5-9de6-4c93-839a-53348ab29c1a · outbound

This paper cites The Hidden Attention of Mamba Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The Hidden Attention of Mamba Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.343979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.343979Z digest=sha256:5f6625d73c0e32186833bfdcaae4db5074a1a35e014ef48bf49ec40b73581dad

Observation e0d11af2-eaeb-46be-b550-ef666c06c5e9 · outbound

This paper cites DeciMamba: Exploring the Length Extrapolation Potential of Mamba.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing DeciMamba: Exploring the Length Extrapolation Potential of Mamba

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.368101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.368101Z digest=sha256:ed737a333369140e8a82d6c04d11f7aecc8abb14ef092d3dd383ce10bd3c4771

Observation b3726c25-d3c1-4271-b87f-f66b23dcd435 · outbound

This paper cites Language models are few-shot learners.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.373579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.373579Z digest=sha256:0b7aef01f061924fdaf4c31112ff5afc2460a0e7c13ed3cf3ac6642929c2728c

Observation 0b53bfbc-b713-4d11-b675-4c8034e50fd4 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.379010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.379010Z digest=sha256:ca5112cebbbd26b2ea455f3f1d7c40ef8ec3f4b15a159ad6636b8edb86538f43

Observation e648a296-7d88-4cd8-9573-f06a8ec462e6 · outbound

This paper cites Rethinking Attention with Performers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Rethinking Attention with Performers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.391034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.391034Z digest=sha256:7d697f706dcd1659ba036c347e134b28fe94378628ced3e0fd77276f05e88caf

Observation ad2c096e-eb0b-4ada-97f1-f0301739c813 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.396665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.396665Z digest=sha256:e330eba70b2e4f1a2ee2d2a7d540289c817bc74648e20843a6be244730f08e22

Observation ff742bf8-2803-48b4-990b-dd784d4ac900 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.407687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.407687Z digest=sha256:1a2782a1556e8172ffc2ba374443fbe945b74bc0fd25c53ecd27703d59ac31e8

Observation 46ff30e7-9068-4de8-a393-2d3f1182e06a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.269547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.413833Z digest=sha256:659c6e810196573d27722e2b0e1c203d2e456c10aa845603ac9a951ee629a66e

Observation 65edb53b-c766-4914-a89f-1d00d1585662 · outbound

This paper cites Were RNNs All We Needed?.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Were RNNs All We Needed?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.425335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.425335Z digest=sha256:b82108a92c7afc0aa4c039d09d892ab4843bc382f895a67bd57d866df32fa2a2

Observation 8701a239-5a12-4284-9d32-4c65459f1f28 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.430458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.430458Z digest=sha256:be672c3a86050f222afe6a83f284fa9f7ce24e3b7de57e8998c9ed77088748ae

Observation f44c4f39-53a2-4934-a11d-6ed4523c579c · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Efficiently Modeling Long Sequences with Structured State Spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.445941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.445941Z digest=sha256:35edb9437f0ce3e6b871115cc99556c2e4ffc57da4c1cc0a96eeb4e5b2aca0e6

Observation 52e4a487-a12e-4691-b1b3-24fed19f0800 · outbound

This paper cites Long short-term memory.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Long short-term memory

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.253392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.450952Z digest=sha256:90271b978862fb97b196d204fa481e06e8c3acc026d85a0c7968bcb998535faa

Observation 196423ba-964a-463b-87b4-7a607dd2f3bc · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.461185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.461185Z digest=sha256:16d865867e75bc64e19f1188a51dd795a95e428c02cdf25fc9cc780c5058cf8f

Observation 7a60ab9c-7c21-4ff1-a378-ee88380691d1 · outbound

This paper cites Mistral 7B.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.466271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.466271Z digest=sha256:6c2be9bb5c0ffb66a9b2ebfffd68203fa21e903bf91945b72187f25c79d2aef9

Observation 4d0cfb62-0c7d-4ea3-8edd-387a9418404a · outbound

This paper cites Scaling Laws for Neural Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.471740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.471740Z digest=sha256:411b7e85b77e341100b027587183f1f562e4f862c8b725ccf8f3fca09f243488

Observation cd04fae6-c460-4fb2-9b81-ae140044975a · outbound

This paper cites Semi-Supervised Classification with Graph Convolutional Networks.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Semi-Supervised Classification with Graph Convolutional Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.477934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.477934Z digest=sha256:7da6b1f08b6dbc9babdd7f6da12752f5ca19901d3b4e6fe4fe5cb2effd81e3d5

Observation 981a46df-05ba-48ed-ba67-ae8ce609bb7c · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Jamba: A Hybrid Transformer-Mamba Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.498531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.498531Z digest=sha256:a8c5ca2e62ae0cb175e85136028354ca049e9318c9113ac93a3acb4f8a4fd644

Observation 604fde47-106f-4e2e-8fe9-8f6ca2dc52de · outbound

This paper cites Longhorn: State Space Models are Amortized Online Learners.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Longhorn: State Space Models are Amortized Online Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.504123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.504123Z digest=sha256:10b64892a449959b3fba2c28b83b640dd1b1911f2d946ea21003017571561c12

Observation 5a25b695-3eb7-42c2-bead-ee61821c716f · outbound

This paper cites Mega: Moving Average Equipped Gated Attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mega: Moving Average Equipped Gated Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.509637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.509637Z digest=sha256:869e3bb538d8a241c2c7841ab94de2855898863a558a4c815991cc4203a59553

Observation 6b830b8c-4773-4ce7-a0a6-2fdf0ffdabd9 · outbound

This paper cites Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.515259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.515259Z digest=sha256:2d78928baf48a03bdfb95afa97d2e1d0b770674b6c533a6c420a83b38e915f01

Observation f7c675f7-ccbc-472d-b38e-bf0e31e13161 · outbound

This paper cites The Illusion of State in State-Space Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The Illusion of State in State-Space Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.520485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.520485Z digest=sha256:f489a0dd9edeaa361c184fa274dec5ccab4881939f91ed16883a06eab4f57368

Observation 5d2451af-a931-47a5-b4f3-c10f01d81e91 · outbound

This paper cites In-context Learning and Induction Heads.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In-context Learning and Induction Heads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.525899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.525899Z digest=sha256:3c53a01b661a7cf10579a9f27c46f05c33fcfd9e5e9360d21a404399b81ce2cb

Observation 94bd2d67-7ceb-4d68-b90d-af485a0d876c · outbound

This paper cites Graph neural networks exponentially lose expressive power for node classification.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Graph neural networks exponentially lose expressive power for node classification

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.237621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.531018Z digest=sha256:617b5efd437fb1164a6c8743fd4eccff485bc3ef60f549a60320e417d0b386a4

Observation bb41ab24-0105-441a-8d79-be64415d5e43 · outbound

This paper cites Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.536068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.536068Z digest=sha256:3d550cd2cd82e6ef24c586c5acb3bce5a75de3aeddb52059e64cdcc9e076bb37

Observation aa7d45c6-9e40-4381-abdd-c0db406053e1 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing RWKV: Reinventing RNNs for the Transformer Era

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.541213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.541213Z digest=sha256:58bfd2e730d719232432fe95e51c705f778351ffdf7be1fe722d1fbd12993661

Observation f1fe9ae7-d4f5-4272-806a-a7f3dccc2fcd · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.546904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.546904Z digest=sha256:f06558a9e85bbfd5553488e9ec6ab6cd7e84d89f2999873d41ff39e783a8d42f

Observation 31794b19-e32f-462e-a7be-4115a891856a · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Ignore Previous Prompt: Attack Techniques For Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.552154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.552154Z digest=sha256:50ebc06d44d7d649a2d15d1f53924e9a8dab6f9748ef46dc249a6846e6ea650b

Observation 3313a84a-12f0-41a4-9f60-523c149c119c · outbound

This paper cites Mechanistic Design and Scaling of Hybrid Architectures.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mechanistic Design and Scaling of Hybrid Architectures

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.557438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.557438Z digest=sha256:e0abe24c20dd273e56351ff5ad5f75f86cd5cacdd7ef60d60e578016764c4cf1

Observation 8b7c8b1c-75c6-4691-824a-824f3bc3daa1 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.563025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.563025Z digest=sha256:355f08ee1c84de58d5d8c64cc91740b90be4aaa0ada2143580fd4ec2336b2287

Observation bb9dc605-5b32-4362-95c2-dfc742b0af19 · outbound

This paper cites HGRN2: Gated Linear RNNs with State Expansion.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing HGRN2: Gated Linear RNNs with State Expansion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.568763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.568763Z digest=sha256:6f5624cf1e9d04804958ccc2ca1de3c79c39ccbeec4e5192b051cadb8167aed1

Observation a2869e78-4e65-4c9a-b094-5fb5bd046626 · outbound

This paper cites Revisiting Over-smoothing in BERT from the Perspective of Graph.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Revisiting Over-smoothing in BERT from the Perspective of Graph

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.574082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.574082Z digest=sha256:9aecc6a3f8b87390456ef481f11de1645af92698f74e1314b7229bd4cdbde2f5

Observation d590677e-9b9b-4c03-97ec-cee95a138271 · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.579902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.579902Z digest=sha256:84cbb939021e6eeb3b7126d5d975770f2ba2cf98b64ee8721d6e364f262c470c

Observation f119f4a3-66a9-40a9-afc9-cff89dbd82a2 · outbound

This paper cites A Length-Extrapolatable Transformer.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing A Length-Extrapolatable Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.584955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.584955Z digest=sha256:635e14cb139bc5a7c159ec69886016cf64dd2dffc432dec328205b253bf024d3

Observation 38c5dc08-a5b4-4e9f-8c95-2a9bc85b24e6 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Retentive Network: A Successor to Transformer for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.589977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.589977Z digest=sha256:68be2e998dc44ccc077132cc33e9026e5387341cfbf5ba9f58ae3b9cb5c14cc3

Observation 7ee84a7d-addb-4452-8dc3-4ffec2706397 · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Long Range Arena: A Benchmark for Efficient Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.595036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.595036Z digest=sha256:ca8944dfc09834c19491023e6afdbdae3d9e7398ae8be47a9c7e9c439d4c3042

Observation 7232c5e4-5021-4556-a012-8dae7549c9aa · outbound

This paper cites Understanding over-squashing and bottlenecks on graphs via curvature.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Understanding over-squashing and bottlenecks on graphs via curvature

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.600245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.600245Z digest=sha256:8916e6df3b1ec2f9a3341eae3ea1b686c6f1856f1f8062e8a3fb287ce722f437

Observation 820c490b-6f6a-4707-8891-06174adc6209 · outbound

This paper cites Transformer Dissection: A Unified Understanding of Transformer's Attention via the Lens of Kernel.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Transformer Dissection: A Unified Understanding of Transformer's Attention via the Lens of Kernel

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.605463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.605463Z digest=sha256:8ae298c19a68d32f7bf8e8e78baa99189ec7edc6aa909b057122e9492ce79677

Observation c0f19b57-13a7-4776-a1a8-7d0ea1f40e1c · outbound

This paper cites The unreasonable effectiveness of the forget gate.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The unreasonable effectiveness of the forget gate

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.610896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.610896Z digest=sha256:1847f834cc850d749844f57755e1e9caee997dad5b2f92abf61187d18bce6144

Observation 3eac356a-996c-4591-981c-374ea6ebbfaa · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing An Empirical Study of Mamba-based Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.615941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.615941Z digest=sha256:0227edf5230984f1205792b78b024a4059b5ac1bcfcebb0a72546afdc9f7abc2

Observation 9fae8463-f14f-42d4-8f5a-1a071cc6734e · outbound

This paper cites Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.621004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.621004Z digest=sha256:7b52d46af1fcb45bd10114ab7a603a0a4666e8676996c530ffd631b3f5a261aa

Observation ff99a6c3-f4c5-4123-957e-d1daa19e3488 · outbound

This paper cites A Non-Asymptotic Analysis of Oversmoothing in Graph Neural Networks.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing A Non-Asymptotic Analysis of Oversmoothing in Graph Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.626180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.626180Z digest=sha256:672a656ec9344ccc3a3555890dd36e8d118ac90c13096a5446fb3fc0a81b92cc

Observation ab356f8d-85e6-4549-bff7-1b23830adf12 · outbound

This paper cites On the Role of Attention Masks and LayerNorm in Transformers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing On the Role of Attention Masks and LayerNorm in Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.631796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.631796Z digest=sha256:f038c01839ccca1b21965b88edfddcdeeddd590bca95be5376f25ed68893661a

Observation 4aa22575-10a5-4c71-99b3-af692f1008c1 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.636905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.636905Z digest=sha256:f174bdcb8870d2b9765e7c0211b9c34bb988bfe7225490f6fb396f79ede7c22f

Observation 4ba87f36-9fb3-4c98-b0dc-8822888cfb0e · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.642183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.642183Z digest=sha256:bc27c8f4c2688f813511996099bab21f5bd5e1b4578bfbebf403369a99c819ce

Observation 0ca922e9-e7f9-4072-98f4-df3502f21d96 · outbound

This paper cites The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.648381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.648381Z digest=sha256:d81a339c0f1e37a10e4c5cad57fc15b2534809f7d5e188fc407c267363d91b96

Observation 7bea6beb-642d-4d91-8fa1-7f41c625fcc8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.653929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.653929Z digest=sha256:4537eea4c265cdcc7827569cf2f1a8dcfbf14a4f6b595c7af2684bb67d847253

Observation efe2d6c0-284e-4b5e-8d57-d9a70d40839b · outbound

This paper cites Arora et al.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Arora et al

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.221479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.659298Z digest=sha256:dfb2af54be0aeb59dcca30126343dd99bd73306a66927e29a41d7dc07fb83c29

Observation cec37677-5009-4c35-9a1f-1d003b909a9d · outbound

This paper cites Linear Attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Linear Attention

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.204856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.664625Z digest=sha256:d21ecf55a1f0a5f6ae8792be80ee5e06eec74a2facf54c6c1cb2242e2d8350bb

Observation 988a1212-ba25-431d-ae6f-15c8d3853831 · outbound

This paper cites Each layer of RetNet consists of a key, 16 Published as a conference paper at ICLR 2025 query, and value transformation, akin to linear attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Each layer of RetNet consists of a key, 16 Published as a conference paper at ICLR 2025 query, and value transformation, akin to linear attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.188224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.669461Z digest=sha256:2edb9608c330532990b3cf98b66986912611ecc5e02334af5df9be0d2ac4c190

Observation ca99d2e5-b3ea-427a-8027-19f3f5abb175 · outbound

This paper cites Similar to Mamba (Gu & Dao, 2023), RetNet shares bt and ct across channels while assigning distinct ∆t for each channel when handling multi-channel inputs.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Similar to Mamba (Gu & Dao, 2023), RetNet shares bt and ct across channels while assigning distinct ∆t for each channel when handling multi-channel inputs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.171024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.674872Z digest=sha256:8ee915cc067af439e9fb7b6c4b7a89deaa1402a14fcf4be5be5c3ab4b8c2ebf4

Observation 8be4f9ad-4423-4eff-b9f8-e7f7cebadf1f · outbound

This paper cites Its computational mechanism can be encompassed by our formulation in Eq.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Its computational mechanism can be encompassed by our formulation in Eq

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.153508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.679907Z digest=sha256:2ccce1d4d0d373104949270d4bb16e8beaa73fe8413cffa1658234cc71a085cd

Observation b2504547-851c-4e66-9a8b-e242a8c6d643 · outbound

This paper cites an unresolved cited work.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:52:35.137753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.684676Z digest=sha256:8382e49ebfef2bb2917db6f37c020fc43e89faeee8a11513cbe3fbcbabb2697c

Observation 9c86dd9e-3253-4af5-a355-f7c35df7859b · outbound

This paper cites an unresolved cited work.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:52:35.120957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.689634Z digest=sha256:b8da0b3738a6a510ec3f73b23c590ebd8ecfa45b5c5b3b88c79742b4728b9cc9

Observation b58c48aa-d56c-4ae0-a6a7-eb3cd43f4f80 · outbound

This paper cites In particular, the dimension of ht in Griffin is equal to the dimension of xt.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In particular, the dimension of ht in Griffin is equal to the dimension of xt

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.099255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.694930Z digest=sha256:ff34cd4503b9dfd29facd9d0cae7e204b8e87ca0c83a34f23325f9c7adfc3ca5

Observation 2c4ae5a1-9287-43b2-be62-a1da2360be1e · outbound

This paper cites This design has quickly become a standard backbone for various SSMs (Gu & Dao, 2023; Beck et al., 2024).

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing This design has quickly become a standard backbone for various SSMs (Gu & Dao, 2023; Beck et al., 2024)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.079906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.700455Z digest=sha256:e8b0a352f464ccdc06a2c99365601914918fce9c5b56b8bda7271b6b75b9dc34

Observation 0ad87dfd-bb5e-4198-97de-4fbe5e2d05d6 · outbound

This paper cites attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing attention

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.062080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.708740Z digest=sha256:35204812d506b44bcf8f281ba8dcfcda8d3abc3872b9fd4b761270c7c476cb69

Observation f362ba20-f132-47cb-9635-ea2101b7fd0b · outbound

This paper cites We consider ϵ >0 small enough, thus, it is sufficient to consider the scenario when |ω| > Amax ≜ maxn∈[N ] |An,n|.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing We consider ϵ >0 small enough, thus, it is sufficient to consider the scenario when |ω| > Amax ≜ maxn∈[N ] |An,n|

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.044358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.714780Z digest=sha256:9a186ba503cfda5e119b95554180cb79b5db320ff5e4a756d5bcf7150a13e5b4

Observation 60b5fc55-623d-445f-9c91-633d71208f2a · outbound

This paper cites Furthermore, let q = 1 − p.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Furthermore, let q = 1 − p

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.027899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.719862Z digest=sha256:f8e1f35a797265986247049ce1feb2e778dd4c4bd86e0cbd59f2c7f87b49a210

Observation 0ff8a946-7cf4-42f8-86cc-c3260a97e1d5 · outbound

This paper cites less than.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing less than

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.010124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.724979Z digest=sha256:1d4109af65b05d9761dcbe301e40a1ed87be6571a5ad28494755bf5b289e1118

Observation e199a764-7053-462f-a8e5-97568a4d0196 · outbound

This paper cites Needle in a Haystack.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Needle in a Haystack

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.875178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.731116Z digest=sha256:1fdd7d1cee2badffca76f253cbc10a009b31e6f7ce54bfdcdb385377e9bb05fb

Observation 92ff088c-6436-4679-92dd-0e678b6714e9 · outbound

This paper cites In SSMs, the class token must be positioned last to aggregate features from the entire sequence.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In SSMs, the class token must be positioned last to aggregate features from the entire sequence

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.858842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.736023Z digest=sha256:2bb1bad37b78a53d51365d2a3a4feed7151bb130b8e7d90c4d09dc594b5ce6c0

Observation 33ea289f-1c1b-45b4-92cf-b1dc038a563e · outbound

This paper cites In addition, our image classification setup differs from Tay et al.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In addition, our image classification setup differs from Tay et al

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.840131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.741482Z digest=sha256:c59444013a08643287eeb05e955526e87b086e2e3148547bbafe98d2d914a6a6

Observation cef29079-19c9-46f9-b0d9-b9f933e1ea1d · outbound

This paper cites The models and training pipelines are built on Arora et al.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The models and training pipelines are built on Arora et al

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.822951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.747493Z digest=sha256:6958f0b24c105fac16eeea508e2ed35ad1e89b3a288a5f228a87d47170429b6e

Observation 688f22a6-b7d7-45c6-a3ba-1edfb924e787 · outbound

This paper cites For each curve in Fig.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing For each curve in Fig

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.803795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.752839Z digest=sha256:3bcbe23c5f9d70b06bdeccbee7038199bc8210bc0e529efd7da5bf5dbb0695c8

Observation 6bf046c4-9441-4621-a89b-01a5b3a1aa9e · outbound

This paper cites The evaluation set is created by holding out a subset of 10M tokens from the training data.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The evaluation set is created by holding out a subset of 10M tokens from the training data

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.781393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.757757Z digest=sha256:1894be360d809ece7077759ced96bdd49eb1270c54f87e6799e5c1ee75e70164

Observation dc07023b-8caf-43bc-ad5c-0828f5bcd519 · outbound

This paper cites We test two block sizes {2048, 8192}.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing We test two block sizes {2048, 8192}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.764751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.762879Z digest=sha256:1fc5533d0aabfef4a56511cfdfe3c35576faafbd18874120825c0c22167079d6

Observation 682063cf-cf7c-4876-a815-9a389b0b523b · outbound

This paper cites 0 A −1000 # , At ≈.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing 0 A −1000 # , At ≈

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.748149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.767881Z digest=sha256:42887680fa993bdd18009cdb533e4230e4d95620e6341f8479109df2149a1230

Observation 55b2cff7-9f66-465f-9b3f-8f1d9719e749 · outbound

This paper cites We consider 1-polarization mitigates locality most significantly, while deepening architecture only relieves recency mildly but deteriorates over-smoothing.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing We consider 1-polarization mitigates locality most significantly, while deepening architecture only relieves recency mildly but deteriorates over-smoothing

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.730788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:52:33.772825Z digest=sha256:0b507c3e010c1939379e72ef4cd8cbd3855fc640ac8ae82618176c0c4700aa5b

Observation c812e534-2070-4e09-9dd2-d6c59daae234 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Training Compute-Optimal Large Language Models

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.455866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.455866Z digest=sha256:83e34b7094cfca66a7e731bf66edbc2c980049b8c5ae858e7d8a566f65ff5ee8

Observation dbbac38c-b5cb-42f7-9537-2af0e4d67ef8 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing DataComp-LM: In search of the next generation of training sets for language models

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.483637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.483637Z digest=sha256:a6dd66769d704617c7823b05bed38b44972d477c384fe3d0e88a1368afd1165c

Observation 88595fe3-5e42-42e3-9ba2-de9ce0d109b8 · outbound

This paper cites On the Properties of Neural Machine Translation: Encoder-Decoder Approaches.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing On the Properties of Neural Machine Translation: Encoder-Decoder Approaches

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.384730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.384730Z digest=sha256:8bcf51fbfb1876074325d50eeb31927dedde318ad402572c1152761afde207dd

Observation 2e6c85fc-6a4a-4d53-8292-32b5122f7358 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.440979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.440979Z digest=sha256:5d0addc2697983a777dabf785e23bbdae961dac912759386a2c9307ca4dd37f3

Observation a86b3051-3238-40c5-b68c-7d4911d25845 · outbound

This paper cites What Makes Convolutional Models Great on Long Sequence Modeling?.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing What Makes Convolutional Models Great on Long Sequence Modeling?

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.489223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.489223Z digest=sha256:00a7a964c9159b86a79f04aeac76da1a5171ee0c505e90085d2caca6ca55d32e

Observation 694c31f0-3289-4e9a-9462-624096c29bfd · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.402204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.402204Z digest=sha256:8b1ab9499a7eb4e65426f2a59b5f60074cc918dd80ba8243a132407f46670877

Observation 02387a9e-f650-4752-9771-413844168932 · outbound

This paper cites Zoology: Measuring and Improving Recall in Efficient Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Zoology: Measuring and Improving Recall in Efficient Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.356788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.356788Z digest=sha256:6caa88aae55fd7f7ba57d0467cbd1aa970319e011ba74493abc2492b34fc2b28

Observation 30770a5f-7dcb-4956-a4ea-d4384b00656f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.419780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.419780Z digest=sha256:330379b402596448860db1ed2dd80f6f176d7db53f2b24ee0b74c6f5196b8731

Observation baced2b1-8fcc-4e09-adb2-d23b25c95fdb · outbound

This paper cites A mathematical perspective on Transformers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing A mathematical perspective on Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.435966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.435966Z digest=sha256:3ce52e436cb5635bcf74db2a203363da3b37592be764f33900a0544654bfecb9

Observation 76ed0eb1-b777-4cee-a63a-b260004c07bf · outbound

This paper cites xLSTM: Extended Long Short-Term Memory.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing xLSTM: Extended Long Short-Term Memory

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.362373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.362373Z digest=sha256:a95e9c38b7c8dd3bf459ccea3dff6a2100217531af28560e580f0d38dacadac3

Observation 64837234-4575-4e0b-b303-0bafb2f189b9 · outbound

This paper cites On the Bottleneck of Graph Neural Networks and its Practical Implications.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing On the Bottleneck of Graph Neural Networks and its Practical Implications

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.351307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.351307Z digest=sha256:87a9e6675cf11c52f23080e5989c11d7df8847d0c13dae9eb84c4f452b611a7a

Pith citing papers

Observation 78136a9d-755a-4e54-a2d9-3a851978ff3c · inbound

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention cites this paper.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.578841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.578841Z digest=sha256:ddc92129b319497e9443fb62ec76796f399f94bae5bb527cdbc510c69042dc7a

Observation e6a00698-189c-4a82-abb3-014f0b22a923 · inbound

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators cites this paper.

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.150123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T01:26:48.026538Z digest=sha256:3590fd11845b0f9c9e52399722b666df1281e53ceb8d69404fae2f99eddeaeca

Observation b59c1c1e-61d6-4d90-a7d9-14f50aeb86fb · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T06:39:25.045565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:39:25.045565Z digest=sha256:60823a5fa0cca09d35134cb2cf6b7eddd5bbd6bde9eab09f22234a47d6701614

Observation 8b9a51b3-bdaf-4644-83ec-dc2a6ca9bd5e · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T02:03:25.442040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:03:25.442040Z digest=sha256:cbb446b529edeee90c6ef3bf1a35268511bec8312a08a210511ab52a4ce5e6f3