Pith. sign in

Paper Citation Record · LEDGER

Customizing the Inductive Biases of Softmax Attention using Structured Matrices

As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2509.07963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07963 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:33:59.310259Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact7
  • verified fuzzy12
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5854e48f-d28f-4313-a9c7-d527029ef4f4 · outbound

This paper cites write newline.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:54.845746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:54.845746Z digest=sha256:1ac0c4fe13ad2b330babae786058933d13565eee316d6ff60067f8f1f3d260d8

Observation 8956a352-369f-4f91-a430-1d8cdc3d8532 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:54.962502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:54.962502Z digest=sha256:625593a49d03abf5bb96c272f6b25a6af20c5d5b515e8f7d893f9220d52e2946

Observation 6784ab5d-1b74-42ca-a359-03d658962c58 · outbound

This paper cites On the Benefits of Rank in Attention Layers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices On the Benefits of Rank in Attention Layers

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.775847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:55.059514Z digest=sha256:4f1c0fb8b944f05f66aca7cc9e63e0c92652337b9113b43c19ab8acd37c6b8a2

Observation f5772a6b-cb04-48f3-b3f9-150f1c575c1f · outbound

This paper cites Chronos: Learning the Language of Time Series.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Chronos: Learning the Language of Time Series

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.131488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.131488Z digest=sha256:2b71f7adb8cfe1ab88c6798eeb227117a58093e400286dfbbe40175da2e02531

Observation 60014f2b-f5af-4e5c-88ee-f179655adeaa · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Simple linear attention language models balance the recall-throughput tradeoff

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.227016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.227016Z digest=sha256:cd798c98ee66582e21b82ef57fe3719b2a8eec7215f543b1b1ba61449f593794

Observation 05fabbd0-b7f5-4520-bf3f-6d0da13df056 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Neural Machine Translation by Jointly Learning to Align and Translate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.310253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.310253Z digest=sha256:55d12cc0367b20e11be4fdc9e328bca33d0f18a395ad75ebe2a8b2c593e5e1a6

Observation 73954db0-52d3-4f4c-8e45-7940f4650f6c · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Titans: Learning to Memorize at Test Time

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.477691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.477691Z digest=sha256:92b71caff0ba0b4db01c7d2c35f56c879fc0cfabab542fc2919d3a4fbe1a78ac

Observation 00c0aeae-9924-4893-8f68-9b83c305ef0d · outbound

This paper cites Longformer: The Long-Document Transformer.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Longformer: The Long-Document Transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.595248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.595248Z digest=sha256:510f362e16c9a00cee216eab450d5ac7c5a3cb715ec98c5c90b53699316af816

Observation 44ab0c7a-f2d4-4491-90ab-0ac5999978e5 · outbound

This paper cites S., Reddi, S., and Kumar, S.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices S., Reddi, S., and Kumar, S

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:03.554091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:55.656699Z digest=sha256:237b22f3cd41cfc9ae6306da47af6606f81f5c7d4b2534c26946c6e3b5d0a2f6

Observation 0d2cb7e8-cf42-4a65-9013-e909c47597cf · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices On the Opportunities and Risks of Foundation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.749150Z digest=sha256:0ad0a9742fe6cd5a1bca5603ff93df4cfa3f2c923f9cdf100c14305149b3b5d5

Observation 6d2ec2ef-be97-485b-99e2-74039c132533 · outbound

This paper cites Scatterbrain: Unifying sparse and low-rank attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Scatterbrain: Unifying sparse and low-rank attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:03.385919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:55.857746Z digest=sha256:0fe233324228c09f56732b4e3b5699db7c7dc3cb96c51dc74e4f6a8089686b91

Observation d9902483-24b1-435e-b2e2-077c5011259f · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.921078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.921078Z digest=sha256:21ed14de3fdf3738e2b5299963ca86571b3735ed8a2ee7923858cdefae818a04

Observation 746228e5-30e1-46e2-847e-e52739af8785 · outbound

This paper cites Rethinking Attention with Performers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Rethinking Attention with Performers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.026504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.026504Z digest=sha256:8a3665425e57743d00342d1192e6028d87738b38e41eb8b9f1b6a38feb93fdb5

Observation f0979cbd-5919-4146-988e-676d9fd70eaf · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.121764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.121764Z digest=sha256:5af558402163d68abf9b4f627333a9b7f62ff78483aea152eccbea16cc753db5

Observation 3cb542d8-6e13-4161-b270-47667b0e4c3f · outbound

This paper cites Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.522778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.219781Z digest=sha256:2f81c8cd28b6e6ef6daa407b5499419c6e1789e18a41a0956e9414c9e5737e54

Observation c37a0a78-bc55-475f-aec6-51f91c26da22 · outbound

This paper cites Monarch: Expressive Structured Matrices for Efficient and Accurate Training.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Monarch: Expressive Structured Matrices for Efficient and Accurate Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.392545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.286090Z digest=sha256:dbda040fa813ffc76700060425edb280ffd5e760a472d30c334d062bedaae6b2

Observation 40dd0845-c4a6-454c-96de-f3dd54359bf1 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.414526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.414526Z digest=sha256:32d4f8e427e17ff42387686bf98833a32c6a6793057d78df78cb77982b463d17

Observation 5decba14-dac5-43e8-8709-73a644e97584 · outbound

This paper cites What Can Transformers Learn In-Context? A Case Study of Simple Function Classes.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices What Can Transformers Learn In-Context? A Case Study of Simple Function Classes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.521170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.521170Z digest=sha256:3e4da745a6de250c2077f3332d013957ab7d55dbe95e98e482a1a49327dc3550

Observation 62a830ed-fc26-4d4b-955a-fbe69e0f2446 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.581715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.581715Z digest=sha256:fd6e04afeea4194768675e956890a030e06492a64d7d3676c3eb7cce9ecf27df

Observation e9bc9332-5f03-4fd9-b5e6-fdca43ca168c · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Efficiently Modeling Long Sequences with Structured State Spaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.648509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.648509Z digest=sha256:2e0bf8a78f547dc4cbd76f69b42e72f336d09ce72cb6c6babdd674ec97be0055

Observation 4decd882-605a-4dbc-8fb2-718693f1c18a · outbound

This paper cites SLT rain: a sparse plus low rank approach for parameter and memory efficient pretraining.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices SLT rain: a sparse plus low rank approach for parameter and memory efficient pretraining

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:03.066582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.741196Z digest=sha256:b44bb661d03819be843fe61644416d78425c7e37b3e6bf23e19407a4631cc615

Observation b5047c79-aa07-4428-a034-848227dae41d · outbound

This paper cites Global context vision transformers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Global context vision transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.760121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.818100Z digest=sha256:5f552eac92ed15408eaabdc4b6081acb09157529199c7e82f5885c5ab41c8efc

Observation d46ffa39-a945-4d74-8fdd-32a069223f72 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.907525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.907525Z digest=sha256:316784138c7d3a53de9967dbceb57440d34442ffd1ec6ccd8e32a6e906596941

Observation 944e5b8a-1d67-4ac6-888c-1474a6ddf9e1 · outbound

This paper cites Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.004385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.004385Z digest=sha256:61ee59ac61ebcfe32aec1ac0b95b0c9500e6f8a2cb029a97e512d5d098f650f5

Observation e85eeb33-3383-455c-9ce0-0c9d6c5ea362 · outbound

This paper cites Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.102946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.102946Z digest=sha256:ad506b55bb0d07e942847cb10ec5d63cb45e592ef505bb8ba12bbb5919df465c

Observation 629fc658-6842-43a5-9268-48bfa93d5168 · outbound

This paper cites M., and Malach, E.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices M., and Malach, E

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.601855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.184572Z digest=sha256:671f28c5c7010396ca1778e2a039540c361474ea25245dc00aef9b93060b63ee

Observation dc923eb5-9062-4833-a75c-873c904a5a86 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.271806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.271806Z digest=sha256:453c2974392ca52f5385c2a3abca3f456099f4a4b7c92c8dfecc190b8e075a3c

Observation 1af3df4b-f13c-49a3-92fa-29353f6f18d7 · outbound

This paper cites Towards Understanding Inductive Bias in Transformers: A View From Infinity.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Towards Understanding Inductive Bias in Transformers: A View From Infinity

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.099758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.360599Z digest=sha256:e0264d6148212d59fa920421e0255a3058ccb37c08451f49569b33c69e2564d5

Observation 889b5a5f-b92e-4b84-9302-b42c2b848916 · outbound

This paper cites Tune: A Research Platform for Distributed Model Selection and Training.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Tune: A Research Platform for Distributed Model Selection and Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.450837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.450837Z digest=sha256:70f498828ac075e70963f4bf2b149f79dd57958f8bb0fd03bcad3d04f15b2b38

Observation ce50c369-7cc8-48ef-b4aa-788e6002551b · outbound

This paper cites Decoupled Weight Decay Regularization.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.535828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.535828Z digest=sha256:c389edc0bc653a521c28201fdfb8aa6536678fc2f7053e7122dbb9e88c9ad4cb

Observation ee85379b-e4ce-4d10-a93b-cb64929f9616 · outbound

This paper cites Non-Vacuous Generalization Bounds for Large Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Non-Vacuous Generalization Bounds for Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.602425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.602425Z digest=sha256:cc84559e85f10857eb4fbc634d4f7142d87954231f08f44c04f70de13706693b

Observation a54bfe07-c46f-47e3-9bfd-17eb2eb1cf9c · outbound

This paper cites Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.675730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.675730Z digest=sha256:e3b7f1dfb10b1654d1ee5c7eb4127453f01c512458476cb6286636516ada9b0d

Observation 64763d31-4fb7-42ac-937a-3a2d2d542b67 · outbound

This paper cites an unresolved cited work.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.729530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.729530Z digest=sha256:5a85f4c0c31bff16fbeeecadf3950d4bee992b500f3513714a5bd36baffdfd2d

Observation 4fa68817-cf3a-4cef-87af-3e41a57ad3e6 · outbound

This paper cites G., Challú, C., Garza, A., Canseco, M.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices G., Challú, C., Garza, A., Canseco, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.399652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.795160Z digest=sha256:a7a028ad032d33bd16882e6a019ac94c56b7960852120960fe73f3d51664e2b3

Observation 5ff92820-96b1-41c4-94b1-e78b02516a1d · outbound

This paper cites Factor fitting, rank allocation, and partitioning in multilevel low rank matrices, 2023.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Factor fitting, rank allocation, and partitioning in multilevel low rank matrices, 2023

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-08-04T21:33:59.912263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.857015Z digest=sha256:b5dab56f926d79d71c3e717cd867fc5d02f5b75f630706de09e9a69a7b5f6601

Observation b5020df4-b203-4893-becd-3f0a8d86f37c · outbound

This paper cites Fitting Multilevel Factor Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Fitting Multilevel Factor Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:33:59.796860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.919742Z digest=sha256:1bc1a166815856f021303582b2ea8e750cf65721856b956a5da75824539c14a5

Observation d6584aec-8a13-4da8-aba1-d60e3818a700 · outbound

This paper cites Hyena Hierarchy: Towards Larger Convolutional Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Hyena Hierarchy: Towards Larger Convolutional Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.984068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.984068Z digest=sha256:ba4c7d6a206e100131b44006f8321d74558f1fe5c1078dc3097ffef57c6eb77b

Observation a911506b-8bb7-4ea1-9a5d-6eee342fa50e · outbound

This paper cites Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.050154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.050154Z digest=sha256:b96145f421d6bf5ac49b84a442e990035822023cd121a0fabf6f1db987a19d42

Observation 6a950f7e-474c-4ecb-9dbe-18605f675785 · outbound

This paper cites Compute Better Spent: Replacing Dense Layers with Structured Matrices.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Compute Better Spent: Replacing Dense Layers with Structured Matrices

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.122029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.122029Z digest=sha256:8c025cbcbedbc9b8d7630c1758973cd0990933c82cfbbc4b3d55b1ed146ae754

Observation 2fbc8239-db77-4e0f-a48b-2c09d1dfc3e1 · outbound

This paper cites Combiner: full attention transformer with sparse computation cost.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Combiner: full attention transformer with sparse computation cost

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.130187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.205239Z digest=sha256:813daec4ed9cf519c957db22cec98e848901a2e56d412af1550c288724c203f6

Observation 585156d6-deff-4d3b-937e-b4d7a7f9fb0f · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Gemma 2: Improving Open Language Models at a Practical Size

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.278888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.278888Z digest=sha256:536608c170a9fc6f532720f4a86c59d57bdd452facb82628ab46b3e2c2d3fa53

Observation cc103613-90e9-416d-bd47-3eae6ba728ea · outbound

This paper cites Representational strengths and limitations of transformers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Representational strengths and limitations of transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.790536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.351859Z digest=sha256:c70af6c8bdf68e2adc787c8bbe5453a249b50e4b21f89499a7178ae6ca1ef14f

Observation efd06eb8-9e63-402a-9676-3b57f7f06270 · outbound

This paper cites A., Choromanski, K.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices A., Choromanski, K

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.583496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.405629Z digest=sha256:8a9eb99edd05d718dfc3883c5f1c90e01a19b5180aa0f643751d63d06573a8e4

Observation 994ef739-8483-4f90-9507-67c99078a52f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.500684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.500684Z digest=sha256:791084846d2d461fa2860fa35863940aff5346fc0916695b2bf731c0ed543dd9

Observation d0201419-cf5a-45f1-a02d-64831d8b799a · outbound

This paper cites T., Gu, A., Dao, T., Rudra, A., and R\' e , C.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices T., Gu, A., Dao, T., Rudra, A., and R\' e , C

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.409880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.560591Z digest=sha256:0a6a185248599d331a31a5184d289f38460ff4c217a72f98a586e0c990d323a0

Observation 7c3565e0-babb-4dd2-ab73-d4025e91785c · outbound

This paper cites N., Kaiser, L.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices N., Kaiser, L

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.625945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.625945Z digest=sha256:fef8826dd7ef194496abf2c99b44dbed4f1d4be3aff94e4e1218813b1ca80bf4

Observation 8849f7da-aa3d-44b9-b700-26f4770a10bd · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.704084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.704084Z digest=sha256:710ff76d71a62a93f0d7eb778ee38ae4196e08c54656027d00afe26d98e916c7

Observation ff859eef-07fb-4a23-933b-1dca443c838c · outbound

This paper cites Building on efficient foundations: Effective training of LLM s with structured feedforward layers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Building on efficient foundations: Effective training of LLM s with structured feedforward layers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.161730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.758143Z digest=sha256:bb71fa331a855f300b1ad51d74aca54563dc718f23c3fcfe1922259a7cf6a8dd

Observation 3033482e-7b63-494d-a4d2-5bb7b2cec55e · outbound

This paper cites MSWA: Refining Local Attention with Multi-ScaleWindow Attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices MSWA: Refining Local Attention with Multi-ScaleWindow Attention

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.820791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.820791Z digest=sha256:9ae3fc4b38b3e5efd18f6b24c664aa1e3c5daad127dcc5d0f35ec3b4546c9345

Observation 0af2bd64-eb77-4913-9748-358a415f6665 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.882869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.882869Z digest=sha256:eb044b9b0a38f6c6f0b5b5a422fd2e6978fd5383dbb810b148b372adcfccaff1

Observation 430a08ec-d21d-4af3-aa1c-1384aa4b3481 · outbound

This paper cites A Spectral Condition for Feature Learning.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices A Spectral Condition for Feature Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.943265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.943265Z digest=sha256:cf036b53ffba0615ed1fcd4007fb8a2e330ff888536d903381280ef8025e7cdf

Observation 203c83a2-4eb1-42a8-aa2b-632ee73ae8cb · outbound

This paper cites BP-Transformer: Modelling Long-Range Context via Binary Partitioning.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices BP-Transformer: Modelling Long-Range Context via Binary Partitioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:59.005872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:59.005872Z digest=sha256:63c741764d716c6d2c211b1c67e1fbda0dfc5419ea582951c4693ac7d99664c0

Observation 1e2b8942-71ca-4d56-bd38-5c4054607990 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:59.087667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:59.087667Z digest=sha256:274941b57281ac000377314bda44a23a3626e50e8c78e275e72aa9a6a65ff7f4

Observation 86486629-e5d7-4f79-a397-9d0df38a749e · outbound

This paper cites The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:00.978183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:59.154952Z digest=sha256:bf15c36418de6d3a21d85a0089ce804bd65f62e619e72400e7abe170c9db9615

Observation 1199fb4c-79df-4cf0-873b-bbd118b4dd8e · outbound

This paper cites Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:59.233909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:59.233909Z digest=sha256:a8d1c52bdd39bb490de0ad8e4b58b5c65d6d2c77f57a320a0f3267b824e30d6d

Observation 53304a06-7977-4aa9-b6d3-f4fb627e8bed · outbound

This paper cites and Soricut, R.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices and Soricut, R

Reference 57

Resolution
verified exact
doi, observed 2026-08-04T21:33:59.482077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-04T21:33:59.310259Z digest=sha256:8ad9ec01a00dde06f793788a18894e6e6ea80b41ca1902bfa2200305cf2f8caf

Pith citing papers

No inbound Pith citation observations are available.