Pith. sign in

Paper Citation Record · LEDGER

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 5 inbound Pith citation observations for arXiv:2505.00315.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00315 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:52:42.980950Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:40.854420Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T23:49:10.742674Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ffe46a2-863b-4131-ab72-d1e2c3f6aa17 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.224119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.643370Z digest=sha256:238850f11eb37048a647da028fb3eeb44f7beffa7d4b461fbb0c20f6ef1aa4de

Observation 674f59e8-ccd2-4108-97ed-e8f0ec5351d5 · outbound

This paper cites Language models are few-shot learners.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.207678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.648989Z digest=sha256:22a295cf7d5a064ebc033ad85664a26b5c4eda55339dc023d353d10d464b6a97

Observation 91dd2db4-7c89-428d-9f6f-e0d65fd878f6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.654726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.654726Z digest=sha256:f0da5967f1f7e02490b34e30f3539b984fce26a82fc07d4e95c734a5eed5ec78

Observation f7a101a1-bf2e-47f8-bfcf-e4ea9790c109 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.660221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.660221Z digest=sha256:db159a4f5a0ab1bf44ad5c86a78b410d6401c8ab8c67f07410525fc1320992a5

Observation 48ce76c7-5857-46aa-b289-193c85a33b0b · outbound

This paper cites The Llama 3 Herd of Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.665341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.665341Z digest=sha256:90a02747c6b398c682232d72dd7b2a3bd03b39e243b8e6929862b51fc7e76360

Observation 8c458f3a-e21d-4839-a237-5d2d2b2276cb · outbound

This paper cites Hippo: Recurrent memory with optimal polynomial projections.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hippo: Recurrent memory with optimal polynomial projections

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.192500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.670227Z digest=sha256:181c5284d5cbc3be3c87b6020cb8208c3877e51d975e6a0a4d0e1593d4863359

Observation 1bb58af2-4c9a-42a6-8bab-d1ab1a90b995 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficiently modeling long sequences with structured state spaces

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.176868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.675684Z digest=sha256:65b8bbdb6aefb848a132b26e53833ec9d368d5fbc956f21b40f2c7ff6d3488fd

Observation 2aac7d60-4c0f-470f-aac9-af63c01643e5 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.680870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.680870Z digest=sha256:cbb4fd0225743ec8cbbe1060c8e2b53a5ac476df46e26e2572cdd17969a569a0

Observation fa12ccc4-ccf4-4ceb-8ffa-a31e03fc3f12 · outbound

This paper cites State Space Model for New-Generation Network Alternative to Transformers: A Survey.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing State Space Model for New-Generation Network Alternative to Transformers: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.685702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.685702Z digest=sha256:96ebe65e58d023dcb083afe3313d1552e43f26be2fb7b3d67429b4a4a75f5c43

Observation ac61648f-4276-46d8-810f-49f57fe562f6 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gated delta networks: Improving mamba2 with delta rule

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.160948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.690539Z digest=sha256:aeb474e3ce13e2f072ee080603c54a3bec20dd1ee767da66979a491e2972ad77

Observation 05ec2dcc-a059-4b42-8a3b-8d6654176d42 · outbound

This paper cites Can mamba learn how to learn? a comparative study on in-context learning tasks.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Can mamba learn how to learn? a comparative study on in-context learning tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.146094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.695560Z digest=sha256:9a69b0f7188b337b9d2345f0c4805d966bb388448a378f60e3fd2862d4ea39cd

Observation 1c99e954-aa90-4c3d-bf08-e9a814bf49d9 · outbound

This paper cites Efficient Long Sequence Modeling via State Space Augmented Transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient Long Sequence Modeling via State Space Augmented Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.700496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.700496Z digest=sha256:1547a041a50538d2ec94bb1e132958fd046481499cb376a95c16a87fd365f52b

Observation 67fb8af9-c2f8-4d10-a259-e0b36483457e · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Jamba: A Hybrid Transformer-Mamba Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.705462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.705462Z digest=sha256:ecd38e623cab6d3b821da63396cbbaf29bb09df6a92424bdd6146dbce95d7954

Observation 8f6f22cd-8b27-4e6c-b2cc-ec3d6af6c1ef · outbound

This paper cites Transformers are RNNs: Fast autoregressive transformers with linear attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Transformers are RNNs: Fast autoregressive transformers with linear attention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.130775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.710076Z digest=sha256:fad943663e7a955e6ba7a85bc4a6ed84b9bd9e94cf646fac9f7d5a709cf0f0d7

Observation ea397617-3235-4149-936b-8cfced97bd80 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Linear transformers are secretly fast weight programmers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.114816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.714879Z digest=sha256:66abfea665e41c82c6c2f3fdfe0ce5f1dcf905f76d5cdb060d69ffb2174c75d7

Observation 5f9e71b4-79a4-4a3e-a4ac-65cc9b65590e · outbound

This paper cites Learning to control fast-weight memories: An alternative to recurrent nets.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Learning to control fast-weight memories: An alternative to recurrent nets

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.099246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.719394Z digest=sha256:8c740e6a9920a9369c6ee4d81329419e22258df6a01e1ed0ece8ce34db76e61a

Observation 1bbffd0b-34fc-4ef1-af68-49f7f5cc809e · outbound

This paper cites The devil in linear transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The devil in linear transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.084166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.724082Z digest=sha256:d19a2e649e0c4ed8893a10c7ecf562b50ddb09c535e58facf0e6843756e1b8e4

Observation 0d691b61-08d4-466c-80ad-36e77e8075ea · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Generating Long Sequences with Sparse Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.728840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.728840Z digest=sha256:fdabc33547ef852a0a37eaa5310d4e53ed0633e7d91269c7101872bf30aa6009

Observation 4cfdd99c-a404-448e-ad95-aab8059a0385 · outbound

This paper cites Big bird: Transformers for longer sequences.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Big bird: Transformers for longer sequences

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.068864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.734113Z digest=sha256:590479d9a7c4b90f54f13f1210d225b0f00ae8c8f73207f64279ee8723fe4891

Observation 7efdb529-dcee-4dbf-8c6d-f1ffbc3b2665 · outbound

This paper cites Longformer: The Long-Document Transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Longformer: The Long-Document Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.739379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.739379Z digest=sha256:98905ef9918701d618b430960ac0b55cdce296e6604f03721e70e15ab787534f

Observation c6d61473-7ce8-4334-a02b-bb4989fdef97 · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Zoology: Measuring and improving recall in efficient language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.053184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.744417Z digest=sha256:2b22f20dbcf749254b7e8576acd34c155e8e3049e2bd842dc6079895eef3c064

Observation addfc2b1-0a86-40d9-b487-4a59362fab1e · outbound

This paper cites Repeat after me: Transformers are better than state space models at copying.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Repeat after me: Transformers are better than state space models at copying

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.037376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.749362Z digest=sha256:2e65bc330d24d06555bd97cf39766c42dfcf9e455e534f260e97e9009333be8e

Observation 97d27218-b36a-437c-a8f8-22d69a02fa89 · outbound

This paper cites Synthesizer: Rethinking self-attention for transformer models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Synthesizer: Rethinking self-attention for transformer models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.018801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.754336Z digest=sha256:911a9977c2b26679a88f72f3382e42d8b5093f7462da086f1d25b81e6e7476b5

Observation ee94e54a-849a-4911-948e-6b3bddeb566a · outbound

This paper cites Fast transformers with clustered attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fast transformers with clustered attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:44.000975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.759015Z digest=sha256:e47523d1784749930b144d7324e5a9650a99a1ef95523864fbb54d71dafc7739

Observation f2c2e1d2-2661-4b26-8489-8bfa366804bb · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient content-based sparse attention with routing transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.984740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.763681Z digest=sha256:0ae03e69c0d4777ff66a90a8bbacc1016773b37dd4c987e2c007eb782e3cfd6f

Observation df8e47a9-ec21-48a9-addb-6d04c5449279 · outbound

This paper cites Convergence properties of the k-means algorithms.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Convergence properties of the k-means algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.968945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.768182Z digest=sha256:e64c5269249e3b057372a0f374b433ebb6ccfcaa55ff371b37994913aef86aee

Observation 8387212e-ab68-433a-b34f-bc2c855ba7bf · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.952806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.773368Z digest=sha256:75970a38dca84f998ac0e3dc714a5c662ce4aaa43c4c2ce76f0280eab24a8776

Observation c6c83f20-df69-4dae-a95b-46562b2bc867 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.936979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.777852Z digest=sha256:d6d5a71975d0725fc6ef716ac596c9f1f780c6d44c4407ba01c38587ff0b23e2

Observation 86b2eedf-af5c-49a7-bc8a-1e6fbca0fb32 · outbound

This paper cites Mixture-of-experts with expert choice routing.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-experts with expert choice routing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.921276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.782593Z digest=sha256:7a8e14ed26638001991599c5825cca8e224026874c150e998636f71c6088371e

Observation 40ceb9a9-6dec-4361-bd20-11a248e593fa · outbound

This paper cites Mixture of attention heads: Selecting attention heads per token.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture of attention heads: Selecting attention heads per token

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.905349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.787352Z digest=sha256:0b933c15dd6e5fb447fb7834cce57b2540ecd4d5ce300acbe56ca681a42ae402

Observation 66d13f37-05af-476e-84c3-2fc369ac3bc5 · outbound

This paper cites Switchhead: Accelerating transformers with mixture-of-experts attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Switchhead: Accelerating transformers with mixture-of-experts attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.889574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.792226Z digest=sha256:5a1f61f39310afaddc98df831a2265b50faf4bf1df489bc9f6f1b1ef1fcf9c86

Observation 404a9814-60c1-4209-abab-8565ce79ca5b · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.873901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.796704Z digest=sha256:81ab20767ab85d7ab77da5f594d1c426f94206bdbb03845bd2ce7df181e4944b

Observation 068ccfb6-5994-4b9f-acd8-dfd75eb1e226 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Snapkv: Llm knows what you are looking for before generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.858773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.801397Z digest=sha256:299c4b8f26c9dcdebd490f35d3f96bb5a7d2b87b67e43298bd77c88aaad3c81b

Observation 8e0e9331-2999-4271-9779-af49e3cf5bc0 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.805826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.805826Z digest=sha256:76c4779e2d8c157b3104401738697e1fb3a363ab33a536c47966a4cf6795f4fe

Observation f2a9a438-53bb-41fc-b1fb-d0332739004b · outbound

This paper cites Gshard: Scaling giant models with condi- tional computation and automatic sharding.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gshard: Scaling giant models with condi- tional computation and automatic sharding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.842793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.810979Z digest=sha256:34b56a3c18ca9833d63cb0aca830a1ffb53c215b291ad494c7495b2a931e8a20

Observation b34db213-1b34-4e7e-9ebe-7e090f1a3c82 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fast Transformer Decoding: One Write-Head is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.815663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.815663Z digest=sha256:2841931936c971ec3aaa4d6c0307da860917f3f3669fb20917dc4be47dd33525

Observation 540e4f27-5d6a-413d-a5e0-29a29372d88d · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.820722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.820722Z digest=sha256:4ee389e0f5644291faad22015a89a54ffa8939b181fa3d320b7c171879aa5f7d

Observation a0f0f4a1-1786-4b40-b9bc-752f5ffb27c0 · outbound

This paper cites Approximating two-layer feedforward networks for efficient transformers.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Approximating two-layer feedforward networks for efficient transformers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.826772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.825493Z digest=sha256:db39f31adbf9a3cc72812bd831e73ec671ae2180eb1f026e60c88ea2ea0ddbec

Observation 9f9b3ef3-6f8a-4be0-8881-7b62eb8a16d8 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.810921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.830748Z digest=sha256:f30be02511e124b5cb51ba9229a2c9a5af02751d40a9e2cc52e3f6224e4c0021

Observation 260485ec-4af5-4deb-bfdf-929c0c6a062f · outbound

This paper cites PyTorch: An imperative style, high-performance deep learning library.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PyTorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.795491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.835539Z digest=sha256:ad1df5d8e3049ebe5620c85d2d406ca3f1cba60a4069908a1755bb5644dd52db

Observation efdf7b8b-469b-4b49-bbf9-c64c0cf4550c · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.840198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.840198Z digest=sha256:7db469451dc9770932ad331b299dbfadc4a2197e7ea86dbf0f535c9c26e26805

Observation 20dc7b7d-80f4-4ead-b1a5-75700f890705 · outbound

This paper cites Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.779647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.845137Z digest=sha256:131c41f4f03eee13b64d27a9ca6768ae04e55f5ab199d135e915202f837de1be

Observation 106e6762-b25e-4913-9f2f-ba02c232126c · outbound

This paper cites Neural machine translation of rare words with subword units.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Neural machine translation of rare words with subword units

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.763986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.849839Z digest=sha256:a9bb67f3258623ea35471c7c4eb889632af3d2e557cce805b9f5f5b04b54429a

Observation 13a8ad77-0d31-4b57-8bb4-189a7a2b2451 · outbound

This paper cites Japanese and korean voice search.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Japanese and korean voice search

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.748580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.854520Z digest=sha256:e9c3704be752b6eba490a86510abd57162546a7b2416c1405c197ceccda4bf06

Observation 7928e4c9-ef5d-4856-8d7c-a4e71a6318b8 · outbound

This paper cites an unresolved cited work.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:43.732176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.858797Z digest=sha256:eeae4467705e97f4c61d6104d0f2865037a47b80c903fc5c4ae3ed828979b648

Observation 52f149fc-d2a9-4d29-b030-70009d6acbb7 · outbound

This paper cites Kingma and Jimmy Ba.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Kingma and Jimmy Ba

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.716134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.863506Z digest=sha256:9632e5bdc2dd45859ce6aab013eda93d18a60c5164566394e0369d21eca9a949

Observation 6927aca1-407f-4441-80c7-bf8056d56456 · outbound

This paper cites Efficient long-range transformers: You need to attend more, but not necessarily at every layer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient long-range transformers: You need to attend more, but not necessarily at every layer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.701133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.868066Z digest=sha256:11cf1e061b7337d54f1a7c33f4db8ca4f874eba62c8a9fca827c26984c32dcb6

Observation 7eaf45c5-b634-4aad-96c7-2ade3ee33bd8 · outbound

This paper cites Efficient streaming language models with attention sinks.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Efficient streaming language models with attention sinks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.686297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.872806Z digest=sha256:19f62b3b0d549c344a3c0a461f96ac208f6cb0978327350fafb18d8c8690ed9b

Observation 8a18061e-fa5b-4bd0-b2a7-b2857453d871 · outbound

This paper cites Reformer: The efficient transformer.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Reformer: The efficient transformer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.670935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.877448Z digest=sha256:6f6d13e4897e1f3ff4cb68d3598f81d339ee59ebe008b79ba36defeadd624736

Observation a68653d7-a2a5-4ebc-9829-364d2e885a1b · outbound

This paper cites From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.882354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.882354Z digest=sha256:b893f2616c405213d3b5e8204c15c65987f67c895c77bf4131c926851ec512dc

Observation 8abcdbb1-12fb-489c-b272-c76849a03c66 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.654662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.887590Z digest=sha256:3f1a1151dbeb77423725b812457f073330c3ce8d2d539a9e55164df01e52d2bf

Observation 926df9c5-1339-4702-85ee-24bd588191ff · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Winogrande: An adversarial winograd schema challenge at scale

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.637170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.892242Z digest=sha256:b1c3c637448dea114fcee06693c75bcab79c23f7099c38f61d919cdd68888539

Observation c8650606-3e83-4abf-9c84-1a20e0a5ab4b · outbound

This paper cites an unresolved cited work.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:52:43.620251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.896965Z digest=sha256:c201fd0a73d52f19b1b245a90cc4c960fa5f383aa1abb57ba49d07b013a6f697

Observation d6316431-8fcc-4494-a085-b5da17a2ab6b · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proc.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hellaswag: Can a machine really finish your sentence? In Proc

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.605098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.901458Z digest=sha256:4faefd6bdc95d535b717a1942b177078ac51b5bd7773cdbc187bfbf0302f979e

Observation fe870209-235a-49c6-a1a7-b8a962d5a49d · outbound

This paper cites PIQA: reasoning about physical commonsense in natural language.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing PIQA: reasoning about physical commonsense in natural language

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.588331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.906082Z digest=sha256:85565cddc2d69c2abd5660e6fa647203c8479c552bdc17a64066060fcb3cd2bc

Observation 3237c19b-e70d-4994-af15-989c465d88ac · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.910654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.910654Z digest=sha256:119538f674afca5036e5ff90193c45120da4a26154cc6db689c8af7cf906cac2

Observation 9c6e41d7-ed50-42bc-9510-20c0b9faba5c · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.915842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.915842Z digest=sha256:891332b7760e12531f8d6458404d86b8585c003529deede7e41a40020c8ba2b4

Observation c682383a-237a-4114-80ae-88bc4e88622b · outbound

This paper cites Mixture-of-experts meets instruction tuning: A winning combination for large language models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-experts meets instruction tuning: A winning combination for large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.571391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.921942Z digest=sha256:c6687adcf5ee1e79c6e73e729ae8061ccea18a213c76082bbfcb45a2d0be2632

Observation 72b5a80f-a5d2-47f2-9670-fa53bd23be89 · outbound

This paper cites Colwell, and Adrian Weller.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Colwell, and Adrian Weller

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.553736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.927034Z digest=sha256:83fc11798497445e16d7db3bc498f6205d3aa6f1e22091351ba3269ba07c4d06

Observation 37c1e652-bb81-454c-8fea-b1eae7554690 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.932977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.932977Z digest=sha256:763b04620c4aa75bed3570894cb40fe63c1f462fc29de1928c2f75f5a9caca3b

Observation af926aaf-0c33-4cb9-a55f-affd6ca805c8 · outbound

This paper cites HashAttention: Semantic Sparsity for Faster Inference.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing HashAttention: Semantic Sparsity for Faster Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.938149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.938149Z digest=sha256:5e09556c267803425082f020d15aa2df5f90048a68cd248523775be8f9ac15b4

Observation 5470def8-36dd-4460-9cb2-799047e85e25 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.943317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.943317Z digest=sha256:2e6d914be82d3a83b797153d822c614abdd7983a2166440608a848e3d6c481aa

Observation 1328b48e-5b56-413a-9dbe-101658e6fea7 · outbound

This paper cites Mixtral of Experts.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixtral of Experts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.947854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.947854Z digest=sha256:92498ea01cdb44a4fd7048ccebd7f9258783ee39688073baff7f81d6902b872f

Observation 3dd3f17d-e420-4b1e-820c-a11159d64c59 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.952805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.952805Z digest=sha256:5b8f0491f9210c3de1f7c7acd20440111df30b5337fbd632a7929da5c0fc1ac7

Observation 2649bc52-0510-4976-843f-38fce4bc8391 · outbound

This paper cites BASE layers: Simplifying training of large, sparse models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing BASE layers: Simplifying training of large, sparse models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.536768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.957515Z digest=sha256:58ecf270e2c15ed070667f710f361e7be28e646bb9f616a054d2f449c74e861d

Observation ddd8c8bf-02a6-4bc5-88ec-d1072d47624e · outbound

This paper cites Hash layers for large sparse models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Hash layers for large sparse models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.519827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.962066Z digest=sha256:dd013be18e8d475708624e71d21845eb7934fab19f46c19dad2fca0589356248

Observation d9ffbcb7-b545-4f82-bb0b-6c49303dc3f0 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.966946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.966946Z digest=sha256:9d71b096dba4755359eb13f02476c42517d0c1d164ab899c0e0494dceb1cee7a

Observation 24d6827b-5f61-4018-91df-4fb99563fb41 · outbound

This paper cites Moh: Multi-head attention as mixture-of-head attention.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Moh: Multi-head attention as mixture-of-head attention

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:42.971702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:52:42.971702Z digest=sha256:e19670b24cadae58ed6495cb7feeb9c541924767e5b682a943fc1a42178f70eb

Observation 66ae2599-b207-4fed-a290-21723eca9b69 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:52:43.503341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.976264Z digest=sha256:7e4f138497dda83f0020a653d6c3b70af0cb60cfec2baa0f3951b0a267b957eb

Observation a5ae2b1d-17f2-4fc4-b392-5bc6efd3458e · outbound

This paper cites On layer normalization in the transformer architecture.

Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing On layer normalization in the transformer architecture

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T04:52:43.486226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:52:42.980950Z digest=sha256:7fc233553210a672ec28cf8baa38c3365b8ce3ef32fa611e0cdfefc1fadebb04

Pith citing papers

Observation 0f65d681-9848-4316-9905-dd2d8b9b35f1 · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.904141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:01fdd50897153889329e41f4cbb1ac9217c0ecec27bd09a163b36ab8e47a5c3c

Observation f4cc3e69-8378-4209-97ef-a5df07deff95 · inbound

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention cites this paper.

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:40.854420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:40.854420Z digest=sha256:f5f5e73741ec01f24fe2f20b02052adfbc7913ca8012735134d3d6799c160a09

Observation 4f23b4b7-3f5f-4dcb-ba2f-797b46b5f675 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:10.746854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:202f40c591aa22705573fe46a4892682b55b429a99c32091cb413de599871372

Observation f25861fd-44c6-427a-8979-dc85ce5d1a07 · inbound

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models cites this paper.

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:42.586782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:42.586782Z digest=sha256:db00f67d3c1a408d9ebf38a0d8c220589929e39f199d78c3864adb42f164edb8

Observation fcfbf863-46b7-49e3-a72a-577cafc0a4c3 · inbound

Compressed Sensing for Capability Localization in Large Language Models cites this paper.

Compressed Sensing for Capability Localization in Large Language Models Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:00:44.103650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:00:44.103650Z digest=sha256:41302a0e36228d06ea195a96416b37e81b4d35ff1367427119353237e12ebdc3