Pith. sign in

Paper Citation Record · LEDGER

Customizing the Inductive Biases of Softmax Attention using Structured Matrices

As of 17 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2509.07963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07963 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:33:59.310259Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact7
  • verified fuzzy12
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5854e48f-d28f-4313-a9c7-d527029ef4f4 · outbound

This paper cites write newline.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:54.845746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:54.845746Z digest=sha256:18eba305b6de6415a3867f55c273a421a44ac24f1d677cd4b2904a0b3441c94c

Observation 8956a352-369f-4f91-a430-1d8cdc3d8532 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:54.962502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:54.962502Z digest=sha256:c9c2a3f357e14415b73e5ccfb94997b90f12a15d6090d818fe3e54b829f653f7

Observation 6784ab5d-1b74-42ca-a359-03d658962c58 · outbound

This paper cites On the Benefits of Rank in Attention Layers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices On the Benefits of Rank in Attention Layers

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.775847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:55.059514Z digest=sha256:d795413333411df789812cada77eedccbf5afec00ab5cb05b7ca83d713e2a0f6

Observation f5772a6b-cb04-48f3-b3f9-150f1c575c1f · outbound

This paper cites Chronos: Learning the Language of Time Series.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Chronos: Learning the Language of Time Series

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.131488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.131488Z digest=sha256:ffe497a5b7ac13a8b0e2f9e15f1a0c28aa61c46d45beb2b5d55b1905d7d9fbf3

Observation 60014f2b-f5af-4e5c-88ee-f179655adeaa · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Simple linear attention language models balance the recall-throughput tradeoff

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.227016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.227016Z digest=sha256:05e6b6959fc400df1a831a651db33c627cd83ec8faadee48c0ea1aaf0230e1f9

Observation 05fabbd0-b7f5-4520-bf3f-6d0da13df056 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Neural Machine Translation by Jointly Learning to Align and Translate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.310253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.310253Z digest=sha256:d3668d89953458320e2393a1b8c8b8b66ac680ae91ae589e06f61c39c89916e2

Observation 73954db0-52d3-4f4c-8e45-7940f4650f6c · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Titans: Learning to Memorize at Test Time

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.477691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.477691Z digest=sha256:0151de088006d3651b815e351baedc644a27e8f5e2fd4d18aa2163c50768968d

Observation 00c0aeae-9924-4893-8f68-9b83c305ef0d · outbound

This paper cites Longformer: The Long-Document Transformer.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Longformer: The Long-Document Transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.595248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.595248Z digest=sha256:75b0e5f86e9d914469a51ce2834e9fcf96de28047dbf3867eea4f0ef9c40c20f

Observation 44ab0c7a-f2d4-4491-90ab-0ac5999978e5 · outbound

This paper cites S., Reddi, S., and Kumar, S.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices S., Reddi, S., and Kumar, S

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:03.554091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:55.656699Z digest=sha256:e4eb7f99bfd93acf67dc8c01f5fc011d0ea2d7d6a1b8350574f4fd1cb4c2edc0

Observation 0d2cb7e8-cf42-4a65-9013-e909c47597cf · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices On the Opportunities and Risks of Foundation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.749150Z digest=sha256:6fac4157fdcffe35a2b09720a2a40df60ef8e7e14d321f835237b501df287d84

Observation 6d2ec2ef-be97-485b-99e2-74039c132533 · outbound

This paper cites Scatterbrain: Unifying sparse and low-rank attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Scatterbrain: Unifying sparse and low-rank attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:03.385919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:55.857746Z digest=sha256:68ccebf14bf0922841b6c4fd284fa40c8f7309161b321b696f1296e6fbb5540b

Observation d9902483-24b1-435e-b2e2-077c5011259f · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:55.921078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:55.921078Z digest=sha256:e9a4bedf1c13d544b00c8e9fd5c9f3f08882b230f858f7e0332e50c6910569f3

Observation 746228e5-30e1-46e2-847e-e52739af8785 · outbound

This paper cites Rethinking Attention with Performers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Rethinking Attention with Performers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.026504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.026504Z digest=sha256:abeecbbff50ef513bb8e29e39c64b1143dc2495a3f25c25b8d55df03390d3438

Observation f0979cbd-5919-4146-988e-676d9fd70eaf · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.121764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.121764Z digest=sha256:df0d6a2030b65ce201f4cde4281e32472f73e7640bf23b6987b00bc2e40b76e3

Observation 3cb542d8-6e13-4161-b270-47667b0e4c3f · outbound

This paper cites Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.522778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.219781Z digest=sha256:0b4bcb08cf020ad17e795e0b2357a7da9c0091edf86e121ab9d3617c3dbddff0

Observation c37a0a78-bc55-475f-aec6-51f91c26da22 · outbound

This paper cites Monarch: Expressive Structured Matrices for Efficient and Accurate Training.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Monarch: Expressive Structured Matrices for Efficient and Accurate Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.392545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.286090Z digest=sha256:c66053515619227ae3e4d5eb341cee57ec9dd0e3688254e052724994b974a09e

Observation 40dd0845-c4a6-454c-96de-f3dd54359bf1 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.414526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.414526Z digest=sha256:20720a61e0fdfdcf8c37921511d2fca29ea1d1df3e2f1aef3601d202051617df

Observation 5decba14-dac5-43e8-8709-73a644e97584 · outbound

This paper cites What Can Transformers Learn In-Context? A Case Study of Simple Function Classes.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices What Can Transformers Learn In-Context? A Case Study of Simple Function Classes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.521170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.521170Z digest=sha256:c45999f4b1ca647df23a6eb7d63b064d808f7c115575ab1d982da80e8d691c62

Observation 62a830ed-fc26-4d4b-955a-fbe69e0f2446 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.581715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.581715Z digest=sha256:933a51a8116cb8ee226f01685211b701fb538d29e19587a9a0f94de4915bf4cb

Observation e9bc9332-5f03-4fd9-b5e6-fdca43ca168c · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Efficiently Modeling Long Sequences with Structured State Spaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.648509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.648509Z digest=sha256:e413e6072c925cebe1148618997ba3edeb306b06d00671175aa31eaaf1fb6720

Observation 4decd882-605a-4dbc-8fb2-718693f1c18a · outbound

This paper cites SLT rain: a sparse plus low rank approach for parameter and memory efficient pretraining.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices SLT rain: a sparse plus low rank approach for parameter and memory efficient pretraining

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:03.066582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.741196Z digest=sha256:4f187ac524d4a546d9c91033481cd3cf36679e90f3a5334faf28c6667f45e6aa

Observation b5047c79-aa07-4428-a034-848227dae41d · outbound

This paper cites Global context vision transformers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Global context vision transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.760121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:56.818100Z digest=sha256:2f660f9619a3872207e5e9cc436c57e145989c753376110dd3aad68a263de065

Observation d46ffa39-a945-4d74-8fdd-32a069223f72 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:56.907525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:56.907525Z digest=sha256:b3ca4ab393da7c3b0316ecab44a2c904992bec08dbcb4daa4ccce0c2ece9f5b5

Observation 944e5b8a-1d67-4ac6-888c-1474a6ddf9e1 · outbound

This paper cites Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.004385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.004385Z digest=sha256:88f574de4f0cd1515b0fbdb3b6c61a980eee38c6b428ed97fc9c4d2d44226d04

Observation e85eeb33-3383-455c-9ce0-0c9d6c5ea362 · outbound

This paper cites Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.102946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.102946Z digest=sha256:acee0531cae33e7e485cc54f224a9334dd49f8d2882166d4725fe5d077f617e8

Observation 629fc658-6842-43a5-9268-48bfa93d5168 · outbound

This paper cites M., and Malach, E.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices M., and Malach, E

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.601855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.184572Z digest=sha256:0d7b0a66991edd03369c1adea28806eae435967e8e9c966e6416ef7f6f56b215

Observation dc923eb5-9062-4833-a75c-873c904a5a86 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.271806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.271806Z digest=sha256:fb2a0c324864b8f80ffd7f34d5b2a84df56c3b0dc7946d79927b882b1030a381

Observation 1af3df4b-f13c-49a3-92fa-29353f6f18d7 · outbound

This paper cites Towards Understanding Inductive Bias in Transformers: A View From Infinity.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Towards Understanding Inductive Bias in Transformers: A View From Infinity

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:34:00.099758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.360599Z digest=sha256:4674fbe88959b1c235817594abef51d60d7674c8fe5056e7a7151f57dafed328

Observation 889b5a5f-b92e-4b84-9302-b42c2b848916 · outbound

This paper cites Tune: A Research Platform for Distributed Model Selection and Training.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Tune: A Research Platform for Distributed Model Selection and Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.450837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.450837Z digest=sha256:e3e50e2a251a5617de5af48010f6ffb2f7d37004e5493e5eb65c8236937a9c29

Observation ce50c369-7cc8-48ef-b4aa-788e6002551b · outbound

This paper cites Decoupled Weight Decay Regularization.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.535828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.535828Z digest=sha256:77a0c6c29742b06868a167b3c3f0909bc6d963f768943ddcf86dcfc20e831b6a

Observation ee85379b-e4ce-4d10-a93b-cb64929f9616 · outbound

This paper cites Non-Vacuous Generalization Bounds for Large Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Non-Vacuous Generalization Bounds for Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.602425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.602425Z digest=sha256:58f967e7c05a121a9da2e5f6e9dc7c4c282c0c6d7cc332dbbe399513836ad08f

Observation a54bfe07-c46f-47e3-9bfd-17eb2eb1cf9c · outbound

This paper cites Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.675730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.675730Z digest=sha256:d092a5052aeb2a70be79320b3de35f9f92162fe5fab0c3cf168af7b6f0bfce52

Observation 64763d31-4fb7-42ac-937a-3a2d2d542b67 · outbound

This paper cites an unresolved cited work.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.729530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.729530Z digest=sha256:57f9114a9cb73f55270c8cf3eea7b37c9cd575f039b93f615a92d347a2b676c5

Observation 4fa68817-cf3a-4cef-87af-3e41a57ad3e6 · outbound

This paper cites G., Challú, C., Garza, A., Canseco, M.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices G., Challú, C., Garza, A., Canseco, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.399652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.795160Z digest=sha256:571542332ca8dd279b0c83a00400338adc2a5e77a00ae340a0771f30932e3c7d

Observation 5ff92820-96b1-41c4-94b1-e78b02516a1d · outbound

This paper cites Factor fitting, rank allocation, and partitioning in multilevel low rank matrices, 2023.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Factor fitting, rank allocation, and partitioning in multilevel low rank matrices, 2023

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-08-04T21:33:59.912263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.857015Z digest=sha256:48eed864a3bf2e1e7504754f9950f501a5a5e9b8a538d0601fa987daec68becb

Observation b5020df4-b203-4893-becd-3f0a8d86f37c · outbound

This paper cites Fitting Multilevel Factor Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Fitting Multilevel Factor Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:33:59.796860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:57.919742Z digest=sha256:d736fda7092b8fd31b677a6db93d14ceee10fbaa50748163c813dd7fe51ba034

Observation d6584aec-8a13-4da8-aba1-d60e3818a700 · outbound

This paper cites Hyena Hierarchy: Towards Larger Convolutional Language Models.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Hyena Hierarchy: Towards Larger Convolutional Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:57.984068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:57.984068Z digest=sha256:df7d12c60fa977626db096e72f3548619eebac42fbb0af1940667c95fd922a1d

Observation a911506b-8bb7-4ea1-9a5d-6eee342fa50e · outbound

This paper cites Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.050154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.050154Z digest=sha256:2666c1ad105da16f54fa47cdf15b8dd9be2eadc3a514243facc3ffb518d8e63f

Observation 6a950f7e-474c-4ecb-9dbe-18605f675785 · outbound

This paper cites Compute Better Spent: Replacing Dense Layers with Structured Matrices.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Compute Better Spent: Replacing Dense Layers with Structured Matrices

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.122029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.122029Z digest=sha256:1c98a209457c53ab0f83d259676881ce0ed88ebc60156eb0fe78b43db37e6fa8

Observation 2fbc8239-db77-4e0f-a48b-2c09d1dfc3e1 · outbound

This paper cites Combiner: full attention transformer with sparse computation cost.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Combiner: full attention transformer with sparse computation cost

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:02.130187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.205239Z digest=sha256:b092ff0581abfb919196e0d3a734f61e4445124958f4c159176bab352928e6ce

Observation 585156d6-deff-4d3b-937e-b4d7a7f9fb0f · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Gemma 2: Improving Open Language Models at a Practical Size

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.278888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.278888Z digest=sha256:23f8fa407da44a33498eebbf0a16ddcaf27619fe3dea21bf81dec8c46a7c2c84

Observation cc103613-90e9-416d-bd47-3eae6ba728ea · outbound

This paper cites Representational strengths and limitations of transformers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Representational strengths and limitations of transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.790536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.351859Z digest=sha256:3662f62c0f9c47b630d5df0fb23a4fe6e25f0c6b5c20138c9f8633b598ffca3c

Observation efd06eb8-9e63-402a-9676-3b57f7f06270 · outbound

This paper cites A., Choromanski, K.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices A., Choromanski, K

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.583496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.405629Z digest=sha256:148de6319e8260ad4d311eeabf80c49b705aa7e81ac46afdb34709efd6b4a89d

Observation 994ef739-8483-4f90-9507-67c99078a52f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.500684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.500684Z digest=sha256:c1f302fe989b93b67473845606484bcf5366a5d3c8bec96489c8eade144dd4b0

Observation d0201419-cf5a-45f1-a02d-64831d8b799a · outbound

This paper cites T., Gu, A., Dao, T., Rudra, A., and R\' e , C.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices T., Gu, A., Dao, T., Rudra, A., and R\' e , C

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.409880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.560591Z digest=sha256:9d1c056e6dcc78aca735c80b505136b267cf4f4b8b0f8de2e62b9998154daa53

Observation 7c3565e0-babb-4dd2-ab73-d4025e91785c · outbound

This paper cites N., Kaiser, L.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices N., Kaiser, L

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.625945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.625945Z digest=sha256:ceb8545e065d59560b12864b15abe9a43c31c95103851f5c7eb872817c1ceab4

Observation 8849f7da-aa3d-44b9-b700-26f4770a10bd · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.704084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.704084Z digest=sha256:b5b153beb6386c0ab64b44045ef2e210f0e3991110ca3de92ad4d9fc54133e50

Observation ff859eef-07fb-4a23-933b-1dca443c838c · outbound

This paper cites Building on efficient foundations: Effective training of LLM s with structured feedforward layers.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Building on efficient foundations: Effective training of LLM s with structured feedforward layers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:01.161730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:58.758143Z digest=sha256:7a3d4eb24fe2c93cee32b1bc2154b3d410e06ecb3428caa077bca8ab8b40e183

Observation 3033482e-7b63-494d-a4d2-5bb7b2cec55e · outbound

This paper cites MSWA: Refining Local Attention with Multi-ScaleWindow Attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices MSWA: Refining Local Attention with Multi-ScaleWindow Attention

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.820791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.820791Z digest=sha256:502ef3cb88854d31be94c403fdebf5094a468279bc1f0b649ef737fd909de7cb

Observation 0af2bd64-eb77-4913-9748-358a415f6665 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.882869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.882869Z digest=sha256:2825784eab9552265a29a53d165189fcc728ad535087871fdb2b8794264f6754

Observation 430a08ec-d21d-4af3-aa1c-1384aa4b3481 · outbound

This paper cites A Spectral Condition for Feature Learning.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices A Spectral Condition for Feature Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.943265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.943265Z digest=sha256:e4f0258c80d2c4d28bac0b9cb9a090feda29d9b119b9af89f94cc2894504edc9

Observation 203c83a2-4eb1-42a8-aa2b-632ee73ae8cb · outbound

This paper cites BP-Transformer: Modelling Long-Range Context via Binary Partitioning.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices BP-Transformer: Modelling Long-Range Context via Binary Partitioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:59.005872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:59.005872Z digest=sha256:1188cbdf725dab965b0cc914c6eb5d1e48fdc3f64244d80da69de72ebc259032

Observation 1e2b8942-71ca-4d56-bd38-5c4054607990 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:59.087667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:59.087667Z digest=sha256:860f245482e930dc0b4ee4ce89fb9689d98b5d5e90df1cc57a7ad49840c8df7b

Observation 86486629-e5d7-4f79-a397-9d0df38a749e · outbound

This paper cites The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:34:00.978183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:59.154952Z digest=sha256:6549064e0e46f3a644dc4d27eab793d12bff41789d11e14fb67bd98984096d67

Observation 1199fb4c-79df-4cf0-873b-bbd118b4dd8e · outbound

This paper cites Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:59.233909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:59.233909Z digest=sha256:23db208a6cf3f9b398d920f5b3146e930dea39ac01b376341af7d95d091bd9d5

Observation 53304a06-7977-4aa9-b6d3-f4fb627e8bed · outbound

This paper cites and Soricut, R.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices and Soricut, R

Reference 57

Resolution
verified exact
doi, observed 2026-08-04T21:33:59.482077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-04T21:33:59.310259Z digest=sha256:fb19fc5194eca5197e4922305767b7def3e9fcc4bb9239336852aea1e69696b5

Pith citing papers

No inbound Pith citation observations are available.