Pith. sign in

Paper Citation Record · LEDGER

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map

As of 13 August 2026, this Paper Citation Record lists 100 of 107 outbound references and 1 inbound Pith citation observation for arXiv:2411.10741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10741 v1

Coverage vector

measured 100 of 107 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:28:30.235477Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:01:50.784615Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T05:01:51.549105Z

Reference resolution

100 of 107 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved55
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca9ffb97-f30e-4a4a-ab97-0253f5b863e9 · outbound

This paper cites Attention is all you need.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Attention is all you need

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.510704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.510704Z digest=sha256:94685b512f13e2eb6a57aae26c2026b386eeee3ecf8e4838c5e4f0ac97b4678c

Observation bcf7d755-9740-4653-9bfd-e574d4f24634 · outbound

This paper cites Language models are few-shot learners.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.551687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.551687Z digest=sha256:a34527acf62b5372bd91d94abe1fb35471d25c17123f0b3a772644b414e9fa5d

Observation 3fec83d1-e2c6-4321-8cd4-0219d1992880 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.558921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.558921Z digest=sha256:66f48dda4d655b3219eda1ecc3b33d37b7bec6458cc77662ab853f5c318daf0a

Observation f304b7d7-e324-49c7-a2fb-d7cdea813434 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.567101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.567101Z digest=sha256:f590ce4daf9fa006abba3d7bbd42945dd05bb696050e211c4b82437052162daf

Observation 411c35a1-b9f4-4ffd-8ddf-6b444874f751 · outbound

This paper cites PaLM 2 Technical Report.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map PaLM 2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.573451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.573451Z digest=sha256:db022ff4dfff5d378e9f4ead1f34a970db359d64d92a17cd007a9851f9efb81c

Observation 0f1cf6f9-d5c9-4f9d-b402-cbed9da12c3c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.579010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.579010Z digest=sha256:317bd3c5db9968b19b479c15c73632f9d6f99bdd3b6013d6221681fd6175cd2b

Observation 81a79af0-9881-4332-a8ef-2afa57beffba · outbound

This paper cites Visual instruction tuning.Advances in Neural Information Processing Systems, 36, 2024.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Visual instruction tuning.Advances in Neural Information Processing Systems, 36, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.585803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.585803Z digest=sha256:91c39b2f79c2c4014804cd64a2099f8e79e331245f804f40cbacdd0dda0bf057

Observation e3bd59c1-1ccb-478f-a8b3-b2cbbb5fc033 · outbound

This paper cites Efficient transformers: A survey.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Efficient transformers: A survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.590723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.590723Z digest=sha256:6a58691637ad792710c43c274ce224ed28d02c1395c625353c37bebee0fb9f2a

Observation e9a4e19b-726e-4d44-8c51-8dc8033bb862 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.595998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.595998Z digest=sha256:5bfd9eb2ecb6a59a86671901a09eb63446dc6cba25489f2fd46d2c8a143161dc

Observation e7d4007c-8d9e-44ba-9931-f2451f752490 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Efficiently modeling long sequences with structured state spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.601596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.601596Z digest=sha256:a70bf2c91bc6164391c0feedb175ef2610915172547a149c23f81c94b6232b30

Observation 12f4dd21-8139-4370-a4f5-9d6c9b69ad92 · outbound

This paper cites Parallelizing linear recurrent neural nets over sequence length.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Parallelizing linear recurrent neural nets over sequence length

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.607247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.607247Z digest=sha256:5b3e75b428c8bfabd6f2eebaff67c6e73aa2f2e0f88f405333f339a9a00e0e07

Observation d523d1d5-c9b6-46de-bf1e-f474972bdc4f · outbound

This paper cites GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.612877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.612877Z digest=sha256:d76fc7fa6bbbeb225defb48021eb5c0f876845c067149e9c86006ffc6212bfdc

Observation 4395c9d1-6da2-4c1d-b620-1cef39dccee8 · outbound

This paper cites The devil in linear transformer.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map The devil in linear transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.618191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.618191Z digest=sha256:e0fb00da4377ab9f7b4b766a64cbcf3b30bc02b2c734cd73d0a936ed467ccdef

Observation a074658b-168d-410a-9370-76f642ca4254 · outbound

This paper cites Transnormerllm: A faster and better large language model with improved transnormer, 2024.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Transnormerllm: A faster and better large language model with improved transnormer, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.622977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.622977Z digest=sha256:5c201545b75c76b2892c0d4ce7538154a69e477fd11d22557e0dbfe501b8cc3f

Observation 4ca69ee4-203d-4ea8-919a-81e66270cd41 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Retentive Network: A Successor to Transformer for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.627795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.627795Z digest=sha256:066222468796f111d25ad967b43df13f79001eb3307fdc97af926205a3963420

Observation 4d2d8aef-48ad-4390-a0eb-b9884e60d086 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Gated linear attention transformers with hardware-efficient training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.633967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.633967Z digest=sha256:a789d99a517f34220879013ca9bafd090759429149d004d7ca3478bb79cb0552

Observation ae5ca70f-14e7-49a7-8c51-aa5cc63cd465 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Mamba: Linear-time sequence modeling with selective state spaces

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.639640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.639640Z digest=sha256:42f24e0a8b14d549cfc553b4a9131299be3a19a24ab365a690e46ace887fe51f

Observation e038d369-6f12-4e59-a501-c4d3fc352517 · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Rwkv: Reinventing rnns for the transformer era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.644996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.644996Z digest=sha256:a8c19a9673adf925c6b22200c46da31f7e2e899d404febf44e6fba5fc00c9acb

Observation 00c89e75-9697-4269-b1d7-89f40bde8b1c · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.652035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.652035Z digest=sha256:661d03bfaa1aebf15d5a4d106f07c3e1f957709217866603de5da60a215c6d5a

Observation 4beb0c1f-cd9b-41d9-9260-6d0c875132ae · outbound

This paper cites Orvieto, S.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Orvieto, S

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.675769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.675769Z digest=sha256:1f0e75d9215816179da60c3ae9cd1f52d48cc86e5d31b88d0c7efb079878e2c5

Observation c4143991-4aef-4b32-8be2-ef5eb2a876bc · outbound

This paper cites Hierarchically gated recurrent neural network for sequence modeling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Hierarchically gated recurrent neural network for sequence modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.735792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.735792Z digest=sha256:d9e0d315844b188a8f349e54763b16312124c4cf1aac2171ae8ee2216b793801

Observation b8b45666-82b2-4cf1-8a45-2b6399e13efb · outbound

This paper cites Fu, Tri Dao, Khaled Kamal Saab, Armin W.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Fu, Tri Dao, Khaled Kamal Saab, Armin W

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.833075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.833075Z digest=sha256:c4733dedd11e25578724439c66ad32a8167e3184d564dd3e0dec1c3018d15c8b

Observation 0a718acb-d305-4257-8e76-2ff52c4a90f4 · outbound

This paper cites Simplified state space layers for sequence modeling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Simplified state space layers for sequence modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.974299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.974299Z digest=sha256:1852a87514d5ce5fe10ccb444037bc6dd33b723d70903481d15dd2f314ba2101

Observation 5df85860-a92d-4c4b-8102-d092493a42c7 · outbound

This paper cites Various lengths, constant speed: Efficient language modeling with lightning attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Various lengths, constant speed: Efficient language modeling with lightning attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.982890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.982890Z digest=sha256:f06b9eaf438ce9f8933c6a928fc1984ca16fe95a50e2e706c0a24cb68243ef4a

Observation a6d3f49f-6ab7-4069-8ce1-f8aa318c24a2 · outbound

This paper cites Modeling Sequences with Structured State Spaces.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Modeling Sequences with Structured State Spaces

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.000348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.000348Z digest=sha256:79d9bd645fea74d347ea4830bd380178410d71d60f625e1452daa80a482347c4

Observation 6422eef5-4427-4174-9cec-ab3fc6db1bed · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map On the parameterization and initialization of diagonal state space models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.008024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.008024Z digest=sha256:03e633816b84351fe45a0b68ee1d7a6793296be58e33076a25638dc3b03a773d

Observation 7fcf0492-4dbe-4afc-aebc-f21aa8d62853 · outbound

This paper cites Diagonal state spaces are as effective as structured state spaces.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Diagonal state spaces are as effective as structured state spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.014127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.014127Z digest=sha256:667a0dd63585ddcfa60c6b3d34e0255e3ea69c9d9d5ee93ce91a20cd459980e2

Observation b49b5146-3acc-4a59-a750-1249905c6c9a · outbound

This paper cites Hippo: Recurrent memory with optimal polynomial projections.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Hippo: Recurrent memory with optimal polynomial projections

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.019317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.019317Z digest=sha256:01b01547c336c58e93c47c5703115f443a2efcb90ee3c67c3a8dbbdd990fbe30

Observation 243f3b77-099a-4da1-a69a-48fcb1d76a2e · outbound

This paper cites Wang, Hui Dai, and Yoav Artzi.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Wang, Hui Dai, and Yoav Artzi

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.025175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.025175Z digest=sha256:3e552fb6c01609dbfc9ab5235e0d32a970055e1c3f0f272bede215f58e029e23

Observation bb9896ed-203d-461e-a0d8-8cecf5fdb597 · outbound

This paper cites HGRN2: Gated linear RNNs with state expansion.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map HGRN2: Gated linear RNNs with state expansion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.098390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.098390Z digest=sha256:0e4c2c1b6670456633a4c14b620770f2febdeedda8ce15dcde0d2f58e86f340a

Observation 6b46549b-4e92-4125-a802-19b45f926849 · outbound

This paper cites Rethinking attention with performers.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Rethinking attention with performers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.205473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.205473Z digest=sha256:32c43638400acb6d0d098d2e1182287687176a1e9f255adc0e24fe77d0cc92a1

Observation ed684719-4680-4e29-b61f-990e27a5fcec · outbound

This paper cites Random feature attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Random feature attention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.243952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.243952Z digest=sha256:05670dcfbcc439458f352a5dc468040439a4f0ad8ddd50d81676a03c78548acc

Observation e525986d-ad46-4c92-9d60-75e0a55514c3 · outbound

This paper cites cosformer: Rethinking softmax in attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map cosformer: Rethinking softmax in attention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.249548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.249548Z digest=sha256:ef93ae9522e884b44f4a2dd22873fe2e71c94387bbc915fdd25fef3e4be90238

Observation 8d7ea96c-729b-4302-9042-60bee9036bcd · outbound

This paper cites Strongly-typed recurrent neural networks.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Strongly-typed recurrent neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.255519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.255519Z digest=sha256:b2eaa1ac6ec07c83637dea85f839449481024f442ca50d30ee571c5f5318fe4a

Observation f80bbdf5-50b2-4e73-916c-1c55c06f98ec · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Mlp-mixer: An all-mlp architecture for vision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.725868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.262440Z digest=sha256:4e010f3a09e989a3b7f7c97000cac139a7bd0a640c7811201dd02ac6929a2e90

Observation 6cf5994b-9113-401e-809a-81b3cd0b3a89 · outbound

This paper cites Metaformer baselines for vision.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Metaformer baselines for vision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.693329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.267457Z digest=sha256:f7dd300afdf6ebc2b9414582dc63f5cb852dc3b715e08d1ebd4dc32520271333

Observation 26b8b53a-c953-40ae-a1e9-3fa3f871e63a · outbound

This paper cites Zoology: Measuring and Improving Recall in Efficient Language Models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Zoology: Measuring and Improving Recall in Efficient Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.273585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.273585Z digest=sha256:464895dfa3bb81c0bbc1433ae90df612b2c1b272e62cb5192aedfdf930dde592

Observation 9b6fe16d-d39e-4104-96a2-5abcbe32b805 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.279038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.279038Z digest=sha256:8008e93ef794a9b632b415ce8b23512b70a0f2ad100ce6cd130236f137b424e0

Observation 4ba3dbde-c758-466c-8df1-b2454eea2b21 · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.664136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.327111Z digest=sha256:33efa24453e5dfa47b2b67bb99c3fffbd9160399c586e70c2495780af85027e1

Observation 35386176-002e-4600-884a-8bfb07ecbd77 · outbound

This paper cites Long range arena: A benchmark for efficient transformers.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Long range arena: A benchmark for efficient transformers

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.631461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.366644Z digest=sha256:bfe793bcf834babad873b0624a72f938aaa0962b0531e253ebf2581426c1f641

Observation f9d915e9-7159-4a45-a21e-2853d70b5445 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Imagenet: A large-scale hierarchical image database

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.410235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.410235Z digest=sha256:4db9c8d51c2c70448c6c0708033c880e720572021221f54d445d7545db4cddb3

Observation 3dd63893-993a-43de-a501-90aeb3c3065f · outbound

This paper cites Mechanistic design and scaling of hybrid architectures.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Mechanistic design and scaling of hybrid architectures

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.575711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.417593Z digest=sha256:3a3c28becac4bf85423b33277f1b5373d752678278211bbf60666e5e3d2f6465

Observation 2c2317fe-ac0e-4a03-99c1-13a36461a170 · outbound

This paper cites Scaling Laws for Linear Complexity Language Models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Scaling Laws for Linear Complexity Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.424463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.424463Z digest=sha256:9f7c7baf93e7e7a4ec1a63ced926a0548013983b931904f224bfef004c2f4ce5

Observation b15e2882-7c0a-4ffe-b4db-74a92faecc8e · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Simple linear attention language models balance the recall-throughput tradeoff

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.431197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.431197Z digest=sha256:a7763b920bf90120d37bae91d0bdbf4bd082f395dd17b696cff339c60ae1adec

Observation cf01d797-e70f-41b7-8bba-ec4192c5ab51 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Pythia: A suite for analyzing large language models across training and scaling

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.536752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.436785Z digest=sha256:c19ee08094c5229bd22f19a42b0c18633c0380e9e66d815d56c2f3df5ee42d03

Observation 8856581a-e5cf-464d-92eb-56221ccd8f07 · outbound

This paper cites GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.505048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.442059Z digest=sha256:683a48f4321a87d7c4c083c04542c9c8895e04b1ed9fd10b6087f2d300848af1

Observation 7f9b955c-0e76-4201-9874-c3c017014bc4 · outbound

This paper cites Toeplitz neural network for sequence modeling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Toeplitz neural network for sequence modeling

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.469887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.447679Z digest=sha256:6b00dfd7bd8aed95c0f52d147c9f39136b75d25b6dbbc304da573d079154c029

Observation e24c270c-9277-4846-bf35-e8d87c1e45dc · outbound

This paper cites Mega: Moving average equipped gated attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Mega: Moving average equipped gated attention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.355946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.453263Z digest=sha256:2fe6ce2ce795d2c826427373600ab5e9e7cd323499a8c0ca33bbcbb322a67d9a

Observation 4424d2b2-fb33-4cbc-957b-729346f9ffe7 · outbound

This paper cites What makes convolu- tional models great on long sequence modeling? In The Eleventh International Conference on Learning Representations, 2022.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map What makes convolu- tional models great on long sequence modeling? In The Eleventh International Conference on Learning Representations, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.268775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.458700Z digest=sha256:6cd0b7ad095a8169a6a90165e217b5453d6c414eee521aa2e9868c37129eeaa9

Observation 629bf5fc-49aa-4138-92b8-34a7aef52f3f · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Training data-efficient image transformers & distillation through attention

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.240675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.535056Z digest=sha256:fd9b487ec8a62ceac55ada280505a515136e34230e131b65ad5fa203267ffa39

Observation 84868c01-1c6b-44a5-89cf-fa258c805042 · outbound

This paper cites The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.211430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.611650Z digest=sha256:60228b0106410bb158f2d39ed44375a95babdd34865d7e684b8bb9afcafac294

Observation 10cf7c8a-56c6-4288-8b2a-3ffb9b64dbc7 · outbound

This paper cites Never train from scratch: Fair comparison of long-sequence models requires data-driven priors.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Never train from scratch: Fair comparison of long-sequence models requires data-driven priors

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.660648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.660648Z digest=sha256:420562cf253bb19550955665bb2a5d9ac97d6a42fd3c95c053b82d5382c1f1b9

Observation 0f6b091f-5f92-4834-9dbf-8dd15516ba39 · outbound

This paper cites Finetuning pretrained transformers into rnns.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Finetuning pretrained transformers into rnns

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.170381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.665711Z digest=sha256:9a3b0c53746f169804c091781dde042b8ecba57073f842c219ec5b428bbb5a37

Observation 506dbc62-6997-4f41-8ccb-916e22b6e0aa · outbound

This paper cites Hybrid random features.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Hybrid random features

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:33.137745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.671040Z digest=sha256:b733b88d1c7b252eafb094d062601a52c7f4f20f8312d2ef0f8f2565a1115a93

Observation ab5508ea-d0d8-4c66-aa9a-392a3dbede42 · outbound

This paper cites Nyströmformer: A nyström-based algorithm for approximating self-attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Nyströmformer: A nyström-based algorithm for approximating self-attention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.676324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.676324Z digest=sha256:be0fa732bf07bb23e98bda3cc892b9fcde4ce7487d6a488f56aca193a650eda1

Observation c6e27a0e-ed54-417b-8b4a-eb023ff0359a · outbound

This paper cites Linear transformers are secretly fast weight programmers.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Linear transformers are secretly fast weight programmers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.680949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.680949Z digest=sha256:d5a037b65c9a924bf9b985fbfaf7aec81f908c4fc80b2ebd0218e659b1ec62c5

Observation ec6caad0-0395-4e61-88bf-a05be0ea85ea · outbound

This paper cites Networks of spiking neurons: the third generation of neural network models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Networks of spiking neurons: the third generation of neural network models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.686157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.686157Z digest=sha256:3c93e593d198b220b28b161fdb0cb696dcea52cbba1c52e1bbb190615ec737db

Observation 5f86ee49-3240-46eb-b72b-32e2879c38f9 · outbound

This paper cites Attention spiking neural networks.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Attention spiking neural networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.691833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.691833Z digest=sha256:15d42cf910fe9205efe7eeb8061074321d95689823b7b4b9d350b5e093eac99f

Observation 95c29368-54d1-4ba3-a17d-11c281a4e046 · outbound

This paper cites Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.696900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.696900Z digest=sha256:065708315a65b3492e2438e68b1236e229b3614db32164142b0bc13fe75a892a

Observation 1aeb4176-1f56-4091-a411-5b82156c1599 · outbound

This paper cites Brain-inspired computing: A systematic survey and future trends.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Brain-inspired computing: A systematic survey and future trends

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.921467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.701911Z digest=sha256:20d75a5001059e95e2de48d1347163aa16d11846ae717940cd46bea353e77fb3

Observation 45aa2243-e590-4d17-9450-abb6016aee68 · outbound

This paper cites Spike- driven transformer.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Spike- driven transformer

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.851746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.707502Z digest=sha256:0edef1320f04d2d8478967d47ba9ccbd0ab15f1ccc1d409e77e0e8c2c7cbf848

Observation 7c796d07-ee9f-4355-a2d0-d1e2f10fe6fa · outbound

This paper cites Spike-driven transformer v2: Meta spiking neural network architecture inspiring the design of next-generation neuromorphic chips.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Spike-driven transformer v2: Meta spiking neural network architecture inspiring the design of next-generation neuromorphic chips

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.783164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.713143Z digest=sha256:157b4e46078ec05945ce670fa83cda10aadf27f47c7b9518adb086be0c438f9d

Observation 7062e33d-3b11-46e4-8155-64bd65e2678b · outbound

This paper cites Transformer quality in linear time.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Transformer quality in linear time

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.703980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.718706Z digest=sha256:25f25d81980418afe29713e1dd0743ceca307d92ca63067a08f5aa6adfa97f12

Observation ecfd5c89-1280-4d5a-bbe7-286b1598b5b6 · outbound

This paper cites Fine-tuning pre-trained transformers into decaying fast weights.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Fine-tuning pre-trained transformers into decaying fast weights

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.666128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.723763Z digest=sha256:15d3e5b211b7652a011c1d03e4a3a535260c6637c85e810cbac695f3a50e8eb6

Observation 161f766f-59aa-42d7-83b3-44d076c0b5f1 · outbound

This paper cites GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.729282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.729282Z digest=sha256:bbe61a0311a15e6b5fef8205d6d89b89b8636a634cf6cb0cde5332c83c5cb056

Observation 2b86f549-cd77-4078-8023-a76d1ef724a6 · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.577892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.823502Z digest=sha256:f811950280fff262351936f961559ffb31305adcbe0e7923d09076ad82083710

Observation 11a28ceb-7edd-4bf9-a9c5-fca84435ff0b · outbound

This paper cites Long range language modeling via gated state spaces.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Long range language modeling via gated state spaces

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.887185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.887185Z digest=sha256:f408580227fd62cd95d4d875eb80a6d67a8f55a06f7116a85d8b5520f31665b7

Observation 00b6d505-2381-4d56-925d-01fc6955dc1a · outbound

This paper cites Pretraining without attention.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Pretraining without attention

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.497604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.894365Z digest=sha256:2de23e77b4b446629107d33331f73a9b88f9f1926a5c6135fb2b0fa6bd13cf18

Observation fe7a2b06-cc62-4fc0-94ba-d319cf7afda1 · outbound

This paper cites Hyena hierarchy: Towards larger convolutional language models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Hyena hierarchy: Towards larger convolutional language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:29.901270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:29.901270Z digest=sha256:6725a243623c38e18d43e103cda867f50637f9a9e0cc1b55f4d6b23845c17fed

Observation c4c2a0be-1caf-4629-90ed-b758caa25858 · outbound

This paper cites S4nd: Modeling images and videos as multidimensional signals with state spaces.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map S4nd: Modeling images and videos as multidimensional signals with state spaces

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.415127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.913499Z digest=sha256:ac4eb2f1d0426d26d19c04e210ba88b1aefa9a08e1f455b4923805ca6ff49b34

Observation 03c088f3-5464-4068-b9db-ed894d9c8f2d · outbound

This paper cites Liquid structural state-space models.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Liquid structural state-space models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.352376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.941277Z digest=sha256:6e991e6931c732ee28cf616f1678f1fca06b6855c28f382b129f62ddb3094548

Observation 4a1b4dfa-a0a0-4314-b565-98af334ad950 · outbound

This paper cites Empirical evaluation of gated recurrent neural networks on sequence modeling.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Empirical evaluation of gated recurrent neural networks on sequence modeling

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.272974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.979063Z digest=sha256:33b465a9fd820141e206264ed58d31dc3ab6eace689f1db934f3e6f0d2f8f38f

Observation a6a0b545-d8b4-44b9-821d-fc5350428c61 · outbound

This paper cites Graves and A.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Graves and A

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.188557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:29.995940Z digest=sha256:f41a57d50c52c45fd8becef162241583f0334ac7acce4de0ea4970014d01b308

Observation 13ea700b-ad21-46f6-ab0e-17dd5785d4b2 · outbound

This paper cites Learning to forget: Continual prediction with lstm.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Learning to forget: Continual prediction with lstm

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.127726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.004082Z digest=sha256:d0de5c33dbde471d5ba9915b1c641447e877d4b6667ec4045d22e322847da412

Observation 9ae0ea9b-a6a7-4814-8fe4-ad346a6eea2b · outbound

This paper cites Linear dynamical systems as a core computational primitive.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Linear dynamical systems as a core computational primitive

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:32.100352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.010505Z digest=sha256:efc48c9b839f3376601cecb39355b1d29c24afcf5ad3d3012b80699257c45e66

Observation b97ea4b9-9621-493a-921d-69a8cc72d283 · outbound

This paper cites Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.017336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.017336Z digest=sha256:5117ac7f1c4cb43d92ba9298e68aee978bee988ef86700a66e96303f7776be4c

Observation 9656a8fd-668a-4800-ada8-3e366323b9da · outbound

This paper cites Quasi-recurrent neural networks.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Quasi-recurrent neural networks

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.997326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.057642Z digest=sha256:fe5fc1fed311da5d79f96c6112af6256a51ef206b481957146f5ee5f55bb0d25

Observation 89904bcf-c5b6-4d08-b38b-4d3a7c3baa6e · outbound

This paper cites An Attention Free Transformer.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map An Attention Free Transformer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.076777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.076777Z digest=sha256:c570d700c94843df4ff6b539ebcc7ba8bc3822979e9d2fddb73dd6c86116ab45

Observation a10137d4-8ff8-4aed-a6f8-8c3d9ac90e6e · outbound

This paper cites Efficient attention via control variates.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Efficient attention via control variates

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.971488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.093858Z digest=sha256:54aead68d953c77597a8e85063c9f17bb331dd9fa4b121a44ddf06e2c58743af

Observation f1cc5f84-c765-4849-a0bd-4f5a0802e4d8 · outbound

This paper cites Fla: A triton-based library for hardware-efficient im- plementations of linear attention mechanism.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Fla: A triton-based library for hardware-efficient im- plementations of linear attention mechanism

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.946743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.100184Z digest=sha256:dbe71d224429c1221d1eb683673b89347e3e9e40b4278770ef28d888b7326567

Observation 3a32fdaf-acfd-4003-bb1f-d01591e68b47 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.106724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.106724Z digest=sha256:43d40bbe2197cf748dd3aa8978924dee7592b6481d5098f90bb4e56bbd2bbcd5

Observation 63fa2bbc-1e30-4a29-82d8-38f5838316e8 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.817955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.126314Z digest=sha256:f33c21250c6bd6e1f14781b0109864a02991bf9c3e10003345450addb0dd0775

Observation abbb2e70-cf63-451f-94af-17053891d733 · outbound

This paper cites The winograd schema challenge.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map The winograd schema challenge

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.795150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.132034Z digest=sha256:fda0d0df76aa1c13ca0e818e31f4c8e125e18e9793c5be8c201b77bdab13ba29

Observation d6c63930-d8bb-4221-bc25-9962c180bbc5 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.137612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.137612Z digest=sha256:ec57c29d7e889a1c2b29c2ab7ed4d035c8c5b6a7417d53db5931b5bf101d4a6a

Observation 5279b2e6-3b2f-4eba-a98f-967999f1c9a4 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Piqa: Reasoning about physical commonsense in natural language

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.763124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.144348Z digest=sha256:724e5b221a3261426849772b137fa985fc5d4d64411f45abd6e28cce8a2ecb34

Observation af29ae7e-d588-4ee3-a1ec-bb34dbc4c893 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.740551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.150791Z digest=sha256:dbb9f07fc99a744ca1ad9e2d7a7e162e5cdbf60dbaf5793e2c39468e77b1be46

Observation efff8d06-af15-498a-ad01-bae114feb876 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Winogrande: An adversarial winograd schema challenge at scale

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.157019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.157019Z digest=sha256:756177602fe3998977a1ab25c6cf1ef752e3984a7040bfd492ebc3a1829e8a99

Observation 7629280a-2d8e-4ee9-824a-1bdc1116f11d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.162921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.162921Z digest=sha256:171fc021d35cfb6fba6c90785248d848041880e952ff0cb52463c1647438f32f

Observation 29f4efde-e11d-4c34-af11-5a56e899fe03 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:30.169129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:30.169129Z digest=sha256:a79811f353ec0eb1b8ae56a337e0493013468925e9e644d170962f58d1cb4dc6

Observation 2c6d6a41-f094-45cc-a756-8e5ddeca5f8d · outbound

This paper cites A framework for few-shot language model evaluation.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map A framework for few-shot language model evaluation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.669207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.175781Z digest=sha256:1ddb2fcc95c61293a13824940c25eb029324535dc9094cdad56875ebabd57556

Observation 060d97d7-5ec3-44c9-b785-a85fda3ff9d2 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359, 2022.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359, 2022

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:28:31.642916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.183942Z digest=sha256:d920bca49e47c43c489e5569cd20e6eb108c75d9b5bf75ec54dbac653b5234c0

Observation 02220aff-416d-41b8-b39b-bd1c065ece27 · outbound

This paper cites including recall, memorization, and compression.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map including recall, memorization, and compression

Reference 92

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:28:31.575640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.190700Z digest=sha256:07e96aa13cd5ae7c4119c23b145e1738d3a07afc644538b9a4c66f41b040c8fd

Observation d0702585-005b-4b63-b85c-569a522ec130 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.479910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.195948Z digest=sha256:40f2ee2508065328ca9e9651c41b1aa902495a69cc830aab35af3efb2eaf76aa

Observation e8d313b8-d128-41b8-80c2-d5501e060859 · outbound

This paper cites Limitations.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Limitations

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.459561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.201146Z digest=sha256:2c62cea015a99b9edd80ab89b1bffc9eccdfe612b347bec68453ec62c6bacb73

Observation e33f776e-cf2c-4ebf-a77f-e5dbd772bab8 · outbound

This paper cites 4 and appendix A2, we provided the definitions and detailed proofs of our theory.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 4 and appendix A2, we provided the definitions and detailed proofs of our theory

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.441915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.206848Z digest=sha256:7072aa9a50f88c9e17793b68d9347fdcac492719b8850a1b281923dc5d69eed7

Observation 9a06316d-7323-497c-a8ce-df87ca9eabdb · outbound

This paper cites 6, we showcased our experimental setup and the results, and in appendix A4, we further elaborated on the details of all experimental setups.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 6, we showcased our experimental setup and the results, and in appendix A4, we further elaborated on the details of all experimental setups

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.423376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.212053Z digest=sha256:1b483296d921cca5cdf033c9848053dd48e18048b43e53a37300a9b0dfb84cc0

Observation 608ea057-ce4e-488b-97a7-45f7797d3b7e · outbound

This paper cites 6 and appendix A4 to ensure the reproducibility of the results.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 6 and appendix A4 to ensure the reproducibility of the results

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.356628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.218861Z digest=sha256:8bb4d4f6814282bcd568e473de149672d8018d446a0c1a52323609f2c0c1ca0a

Observation f6550875-6b7e-496e-9cbe-5012a707947e · outbound

This paper cites 6 and ap- pendix A4.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 6 and ap- pendix A4

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.287000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.225249Z digest=sha256:2b34f59b6ec6567f60acf5cf00b3fd7c79d8670c9c6f8a47800b2d943cf0d7e2

Observation 2efa4890-71c8-4c0e-a45e-3d930f84f314 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Guidelines: • The answer NA means that the paper does not include experiments

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.187468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.230186Z digest=sha256:7d49a3bff7897559f064f92f091399ebcc9b46ebfcdbe2a86b13b1a42ae471a5

Observation 346dfad7-c842-4997-96a9-6c1b695e912b · outbound

This paper cites 6 and appendix A4, we provided details about the computer resources used for the experiments.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 6 and appendix A4, we provided details about the computer resources used for the experiments

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:28:31.095865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:28:30.235477Z digest=sha256:c63b1f2f8acc4d19851213e74e6412589bfd635e4f77595575118bc72d671c36

Pith citing papers

Observation eeeb9e82-af74-43df-a15d-25f3b9ee0a81 · inbound

VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis cites this paper.

VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:01:51.556459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:01:50.784615Z digest=sha256:d51b4e079b7c2cae22beb6c4eff91b6d75cf3246161aff382a7cd53ef2e01e34