Pith. sign in

Paper Citation Record · LEDGER

Accelerating Attention with Basis Decomposition

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2510.01718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01718 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:54:36.489014Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73142ff2-4f01-4393-9534-36de299848ad · outbound

This paper cites Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman.

Accelerating Attention with Basis Decomposition Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.229619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.229619Z digest=sha256:48f65f28ea937664a78fdbbce18c6cadfdb37956ffa92f79a5a98ce3c2a95515

Observation 640eeb3a-151f-454f-9f2f-4da6f2fa88f4 · outbound

This paper cites Longformer: The Long-Document Transformer.

Accelerating Attention with Basis Decomposition Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.269838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.269838Z digest=sha256:b8d3ac28e95fa9876820a4a5744818d86b87bb88ff70e98e688ecc534ff5fde2

Observation a7045ed7-be93-4994-92ba-59ddb6e45bc7 · outbound

This paper cites Language Models are Few-Shot Learners.

Accelerating Attention with Basis Decomposition Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.322033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.322033Z digest=sha256:ad06d1e2eca2d210b134855609a5bc98d5e32645ba441b6f2b956fe13a78e7ec

Observation 00d8f182-ab52-47d9-8058-734c395a446f · outbound

This paper cites Linear least squares solutions by householder transformations.

Accelerating Attention with Basis Decomposition Linear least squares solutions by householder transformations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.370688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.370688Z digest=sha256:27d5d49201c4738409371e4b65a039fb7604e90bdcb0169fa96cbcf7167506ed

Observation 5a81a44c-5501-4bb1-a19d-95d6a2fb30ee · outbound

This paper cites u ker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Marcello Federico, Sebastian St \.

Accelerating Attention with Basis Decomposition u ker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Marcello Federico, Sebastian St \

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.457634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.457634Z digest=sha256:33a8e5b6eacef52cb6381ed014c97fc7efae3d9d38018d6e8fcc4727c98d98df

Observation 5af29c4f-d56f-4c78-ac2c-3a8c2f8f77cb · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Accelerating Attention with Basis Decomposition Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.547368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.547368Z digest=sha256:61f75cbd448ee7d0b2214f4cacbcea7eb1f8186f5354c039fa4288b2824656ba

Observation cce563b0-900f-4561-a90d-b07ee1fdde09 · outbound

This paper cites Rethinking attention with performers.

Accelerating Attention with Basis Decomposition Rethinking attention with performers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.602091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.602091Z digest=sha256:4a1dc8d0b0eddc04b73d5717e293ff45027fc2817b6893bd71484d9353428e83

Observation 20be47fb-d94d-4b85-8eb5-9b30cf963f76 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Accelerating Attention with Basis Decomposition FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.663535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.663535Z digest=sha256:045fc1c6c134b218780e8d7e0d2ffc731ba6230e273322f18303e76db3afe272

Observation 21a61534-ad57-4eeb-a424-c151aadd6afb · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Accelerating Attention with Basis Decomposition Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.706584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.706584Z digest=sha256:58fdca956ea1a709c21f03ea4a9a9e9a0231aeb9a2775b5541ae5003010f7bec

Observation b6941ade-a66e-46b3-b115-1816909e2f36 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Accelerating Attention with Basis Decomposition An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.763214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.763214Z digest=sha256:f579fdb6ec6d73a25ed0a62055bd20eabff1a519a4ee9c284a07cc7bbd9a11a1

Observation a2bba1f1-1e01-4127-83a6-4f2bc1d452d3 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Accelerating Attention with Basis Decomposition Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.802792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.802792Z digest=sha256:80931c3187c5f86e1461e4fa01f0967ffc7d8e03289df5b0d4b8633d51602055

Observation 6010a440-46a5-4e74-89ff-412c8e4fa407 · outbound

This paper cites OPTQ : Accurate quantization for generative pre-trained transformers.

Accelerating Attention with Basis Decomposition OPTQ : Accurate quantization for generative pre-trained transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.824068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.824068Z digest=sha256:303638a48f5d59d0fad0c2de991dd5f3c4b25bca97df506e8d07e73c7555c6ad

Observation 5a9db1f6-3352-4f02-9ef6-f5ca335451d4 · outbound

This paper cites Strategies for applying low rank decomposition to transformer-based models.

Accelerating Attention with Basis Decomposition Strategies for applying low rank decomposition to transformer-based models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.865017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.865017Z digest=sha256:0ac5bcfeb63aa4d7930e65694983bf5078ac3a5e3f7e1e3c9613869d1973d734

Observation a057aba3-d32e-4e15-9bf0-b17222eaee1d · outbound

This paper cites SLTrain : a sparse plus low-rank approach for parameter and memory efficient pretraining.

Accelerating Attention with Basis Decomposition SLTrain : a sparse plus low-rank approach for parameter and memory efficient pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.910134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.910134Z digest=sha256:7b55c50ad4b380f3dcc3c50231d5c518aca31040c2a8bdf676e44cec089d7b0d

Observation be4e18b6-074b-42af-9377-38971a8d2709 · outbound

This paper cites Language model compression with weighted low-rank factorization.

Accelerating Attention with Basis Decomposition Language model compression with weighted low-rank factorization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.931321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.931321Z digest=sha256:38025be40122873a10674cc5101fa5413c59c06d3828d58aae4c5b0e268af0d3

Observation 2a63d98f-6da0-464d-9196-37cc203100a1 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Accelerating Attention with Basis Decomposition Lo RA : Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.029756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.029756Z digest=sha256:472c2b7d206f9a1606fb10b591190b97f7acf6b10a55f5f5b97c798133cc55cb

Observation 2b676826-9b17-4b04-b090-538dc0d1f1d4 · outbound

This paper cites From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications.

Accelerating Attention with Basis Decomposition From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.091978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.091978Z digest=sha256:9e2ec3415c32f3c16750e0face2f55b8f737130fabd8bc96b094c8f826bf9881

Observation a6a14013-6408-4d60-858d-e52a7e75f9e7 · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

Accelerating Attention with Basis Decomposition Exploring Low Rank Training of Deep Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.209617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.209617Z digest=sha256:c612f05d3f76964d7f1ca8bc76526cb97c917833efb534375b2fc6361f12cbe0

Observation 631783ef-d799-40a2-8984-3a39a3333220 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Accelerating Attention with Basis Decomposition Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.333202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.333202Z digest=sha256:55dcf156a7524283acd0898908580360f2ad8a3c235ce917072a7ec1a704a5e9

Observation dedb3b3c-0143-447c-86ac-5a807bdcfb2e · outbound

This paper cites LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression.

Accelerating Attention with Basis Decomposition LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.437126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.437126Z digest=sha256:1c2066fb56704cf2a42970e54fdf1eacf4bb3d572cf609be26bf75087ca0d17e

Observation ec62ac58-2159-4e00-978d-e4bd905ef6e1 · outbound

This paper cites Tenenholtz, Lester Mackey, and Nicolo Fusi.

Accelerating Attention with Basis Decomposition Tenenholtz, Lester Mackey, and Nicolo Fusi

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.557629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.557629Z digest=sha256:3d97cef05a9c481c552f0041ea7cc98e1422f3ff14622a303f44773e53dd0132

Observation 0f83de22-e96d-4476-9f62-7053a68070a3 · outbound

This paper cites Reformer: The efficient transformer.

Accelerating Attention with Basis Decomposition Reformer: The efficient transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.651506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.651506Z digest=sha256:5e710ffff9060efcf9ce09ea095fa00fc01b3c3deaa699c5dc74ffdbb74aa466

Observation 31db1703-5b8b-4a93-b2ca-a68b152ecc44 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Accelerating Attention with Basis Decomposition Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.752334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.752334Z digest=sha256:8c24bc492925c4e505224ca6ec12cd319d203512e6dd329b2f37a60a6b5c77ee

Observation 643ba7d0-1136-467c-b008-08c2923f039d · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Accelerating Attention with Basis Decomposition Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.894154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.894154Z digest=sha256:eb1b08594cffd5b1e36fd6d1489099de431cfe46e26a6db6c9daa77904fa3025

Observation 01fa899c-3f53-4895-be6e-7047fb963ed7 · outbound

This paper cites L o S parse: Structured compression of large language models based on low-rank and sparse approximation.

Accelerating Attention with Basis Decomposition L o S parse: Structured compression of large language models based on low-rank and sparse approximation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.978012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.978012Z digest=sha256:c7ff939dbfda53c574e52f1890c7257f2d4b551285ef15e47665585bd5f8c638

Observation f261d3aa-14a0-4a8d-9c5a-d529ae8ba1dc · outbound

This paper cites Relo RA : High-rank training through low-rank updates.

Accelerating Attention with Basis Decomposition Relo RA : High-rank training through low-rank updates

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.133004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.133004Z digest=sha256:e0ccd53a1ba0d9e3c488f197103f68802c38e3a1e1f3680cf05ed8f9fa721c55

Observation 77cdf235-0bfd-4f03-99a5-b557d1825d12 · outbound

This paper cites MoDeGPT: Modular Decomposition for Large Language Model Compression.

Accelerating Attention with Basis Decomposition MoDeGPT: Modular Decomposition for Large Language Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.247901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.247901Z digest=sha256:667758c5b27cc1f358dc51b8b52319f98586cfe9a6327c9bac5e0aff8e9ad1f7

Observation 83137743-3cb1-4ff3-b11a-6cf1c74487d4 · outbound

This paper cites Duquant: Distributing outliers via dual transformation makes stronger quantized LLM s.

Accelerating Attention with Basis Decomposition Duquant: Distributing outliers via dual transformation makes stronger quantized LLM s

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.347427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.347427Z digest=sha256:0b42616a5f104c006c621cab2a9d6e2b6ede3c1576cdf903e1de1189bb9ba946

Observation 9f0d11ae-4f6e-4531-a7cc-a1a6c23fc160 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Accelerating Attention with Basis Decomposition Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.471274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.471274Z digest=sha256:a512c19bb449c5555157d197478a2a7d74d4f41dc1cb16973b7c94870a2189b9

Observation 5ae5674c-ffb0-4b20-839e-613f321f3d60 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Accelerating Attention with Basis Decomposition DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.527232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.527232Z digest=sha256:397a94775fdd139794029dd32697c6903b578f7ac2b2eee0414f9850490a7b93

Observation 07b3b64d-3b08-49e8-8f1c-d43f6e3b80c1 · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating Attention with Basis Decomposition DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.615317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.615317Z digest=sha256:99b61285c07a621ccdacc1c64a9649b8d5c0778c9a56d6ae1dc165db6e686256

Observation 0258c38e-9abe-4a8a-b13e-c7ebc7dac914 · outbound

This paper cites Dora: weight-decomposed low-rank adaptation.

Accelerating Attention with Basis Decomposition Dora: weight-decomposed low-rank adaptation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.698827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.698827Z digest=sha256:22c839f84edec572e3238bcaf690570c07e257f00157b7cc7458078cae1917cc

Observation dd5384e6-bb56-4bcc-8dd3-d47a66c0135b · outbound

This paper cites Eora: Fine-tuning-free compensation for compressed llm with eigenspace low-rank approximation, 2025.

Accelerating Attention with Basis Decomposition Eora: Fine-tuning-free compensation for compressed llm with eigenspace low-rank approximation, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.861188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.861188Z digest=sha256:a3375afb476e1a0b8a1e9dfb2e3b1a0678abe98d1b4228279a51b8929e7596e9

Observation 7ac60136-ac83-4676-a145-dda4bcac9790 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

Accelerating Attention with Basis Decomposition Llm-pruner: On the structural pruning of large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.977760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.977760Z digest=sha256:98251d8d451fe32f483071f41783aeb80f8df616a2ff6d38f292b59056fd237c

Observation 953b5e5b-d8ec-4d75-a517-f86fd2b5d80d · outbound

This paper cites Pi SSA : Principal singular values and singular vectors adaptation of large language models.

Accelerating Attention with Basis Decomposition Pi SSA : Principal singular values and singular vectors adaptation of large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.097917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.097917Z digest=sha256:0aeb4613d271051d19091a796973076a7d943fdb1aa20efd8495cde710d1b7d5

Observation 3dbea6ef-0b63-4017-9997-9c4ea74232fe · outbound

This paper cites Accelerating Sparse Deep Neural Networks.

Accelerating Attention with Basis Decomposition Accelerating Sparse Deep Neural Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.214208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.214208Z digest=sha256:ba4d83de2d74a76d04083a0616bd9b2020b38af9f19e526b9dac0997b3b1b102

Observation d439b2ad-6e08-4225-bfbd-55466bb6f4e4 · outbound

This paper cites Dobi-svd: Differentiable svd for llm compression and some new perspectives.

Accelerating Attention with Basis Decomposition Dobi-svd: Differentiable svd for llm compression and some new perspectives

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.322566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.322566Z digest=sha256:e162ef96e52e9b76f9e8cdb85bc5963e7cd29c0776b3f994599d01179dc4383d

Observation 6f961544-b9b2-4dd2-8839-cc51cc6df62f · outbound

This paper cites Improving language understanding by generative pre-training.

Accelerating Attention with Basis Decomposition Improving language understanding by generative pre-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.440915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.440915Z digest=sha256:bb4b5eb805266df2f589c88edc12362228f616ab1de2006102d958ccb6b8e47e

Observation c43b9212-5043-4bf1-9d40-21ba7e391419 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Accelerating Attention with Basis Decomposition Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.536976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.536976Z digest=sha256:20dfb5bec1cb5f6a25de97617b91f18070595ba532905109b8f6cbad06bde199

Observation 278dc369-fde2-4308-9668-e49d0114f923 · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.587104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.587104Z digest=sha256:6d7bcbd22b60bd3923e753672d19c70d3eb393f6ab0e61fd1e662c847c37e9ae

Observation 31e01949-78d7-4159-beb9-0126df88ca27 · outbound

This paper cites Compressing large language models using low rank and low precision decomposition.

Accelerating Attention with Basis Decomposition Compressing large language models using low rank and low precision decomposition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.690785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.690785Z digest=sha256:ea9b0f5eda70fcf33ff9482bb83812d52d3c5c4944d72aac1077e00d6ac69478

Observation 342da8c0-a4e9-4b3f-9574-f2d413e0f6b4 · outbound

This paper cites ESPACE : Dimensionality reduction of activations for model compression.

Accelerating Attention with Basis Decomposition ESPACE : Dimensionality reduction of activations for model compression

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.809317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.809317Z digest=sha256:373485add83d231c6c677d402e0c5bae9947fa595686674f59f06034abb29edd

Observation e920ecbc-9788-49b6-8e95-0174b00261b3 · outbound

This paper cites Robust low-rank training via approximate orthonormal constraints.

Accelerating Attention with Basis Decomposition Robust low-rank training via approximate orthonormal constraints

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.919849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.919849Z digest=sha256:4eb616dab712263789fcf469a1d344b669446574e46012c6206a86a4a942c395

Observation 1ea1abe1-5963-4ead-a7f4-2d53663b6ac1 · outbound

This paper cites Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations.

Accelerating Attention with Basis Decomposition Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.055161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.055161Z digest=sha256:8c599b411b5ecb6fdc31091f24171422acec1a4369ccda91fdd48879c37f9d54

Observation 93e826c3-8a4c-4eb0-a545-a663f681d28b · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

Accelerating Attention with Basis Decomposition Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.234225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.234225Z digest=sha256:a77986e41a1519e972df3396925806df79d42e3d9f91b5899d1f2c34cf85f522

Observation 586d5a36-5a80-422b-b19e-72f3bccac425 · outbound

This paper cites The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction.

Accelerating Attention with Basis Decomposition The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.403786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.403786Z digest=sha256:0e3b0fa5c9e76410ee2e20c0fa6f0eaa1dace8c11ad28a67dc83cdc5ebe65003

Observation 528f9003-faa7-48b4-80d5-90af2d766239 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2021.

Accelerating Attention with Basis Decomposition Roformer: Enhanced transformer with rotary position embedding, 2021

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.560871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.560871Z digest=sha256:eb5a5a622e7c0f8fb61bebd814aa699a0622dacee189760d43c5834907e103fd

Observation 18fd2129-8fbe-4609-949c-f15b7ecfa635 · outbound

This paper cites A simple and effective pruning approach for large language models.

Accelerating Attention with Basis Decomposition A simple and effective pruning approach for large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.725960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.725960Z digest=sha256:b66d48d690af66b0529eae9a5f326b61408a8b93649a91f9696229dc430f1ba9

Observation 3ae9a8c8-8529-4874-9453-f8d373e940e4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating Attention with Basis Decomposition LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.850727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.850727Z digest=sha256:98616438590a7488fa7b3a04dfdcc71360b49d5e3c52e97663727d913115987c

Observation 3e2ec9d2-3863-4724-a5b5-713fdfd1aa2f · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.978563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.978563Z digest=sha256:b2802c113224fb4b2f06639f53ae431a571574eb4dc0e156f4191ef935334e29

Observation f0136548-5b07-446f-8cd5-045e713c7f8b · outbound

This paper cites Attention is all you need.

Accelerating Attention with Basis Decomposition Attention is all you need

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.063112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.063112Z digest=sha256:a41374c65d9b9793a27838aeee9a6e511ddd43dcb4f5a38fbec8553b9d07f390

Observation bec3bf8d-1d0f-4873-9ac4-a9fea3897be1 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Accelerating Attention with Basis Decomposition Linformer: Self-Attention with Linear Complexity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.148840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.148840Z digest=sha256:cab6e401d801febe1e2a16403c05c165516a9621adea280fa0ed311003112a21

Observation d362884c-1a04-4ac8-93cd-c40cc7ca3094 · outbound

This paper cites SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression.

Accelerating Attention with Basis Decomposition SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.258159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.258159Z digest=sha256:d880ce66cd65cef62b4c8840b4f85c885fd96f9a7efb74f8ab0bf7d82f409711

Observation c321b518-a77c-447f-97dd-88711359f07e · outbound

This paper cites S mooth Q uant: Accurate and efficient post-training quantization for large language models.

Accelerating Attention with Basis Decomposition S mooth Q uant: Accurate and efficient post-training quantization for large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.361956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.361956Z digest=sha256:a9c5cfa1af102d91160c33091d6d2ac043ee5353a380e690328de4bfa3359b91

Observation 3a849161-2923-4d55-bd1f-e3a7efe8544a · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

Accelerating Attention with Basis Decomposition ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.464572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.464572Z digest=sha256:2cdc6a0ae17039de084f7482bed0da2443bc525ecb80dfcefa366cedf44292c2

Observation fd338049-9301-4ef7-b7ea-92b6c9f09adf · outbound

This paper cites IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning.

Accelerating Attention with Basis Decomposition IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.581040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.581040Z digest=sha256:ffcba430f934cc14de9fb2e7cca9cdc23570c9fee885a091fc088928c7897802

Observation a45d9535-e250-4bd4-bb31-6d366b7a860a · outbound

This paper cites Adaptive budget allocation for parameter-efficient fine-tuning.

Accelerating Attention with Basis Decomposition Adaptive budget allocation for parameter-efficient fine-tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.660902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.660902Z digest=sha256:db26c82fff5a3f9752f705a156889268e660cea80aa8dd4c1a1b27cbfa4a70a0

Observation b1bce96b-694c-43aa-9b5e-4f059156fc26 · outbound

This paper cites OATS : Outlier-aware pruning through sparse and low rank decomposition.

Accelerating Attention with Basis Decomposition OATS : Outlier-aware pruning through sparse and low rank decomposition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.775920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.775920Z digest=sha256:439fc8e09f1b5463ed6f5a47c56ef290c737e4d82ee77b3a2a8526eeeaaca893

Observation dc3d7516-ad5f-4dd5-a41c-de003089dfb9 · outbound

This paper cites Plug-and-play: An efficient post-training pruning method for large language models.

Accelerating Attention with Basis Decomposition Plug-and-play: An efficient post-training pruning method for large language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.900097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.900097Z digest=sha256:0ac983f0534687052ca67b7c690f4c44cfcde3375a3b7ae3c0c3d2e969b9b9bf

Observation 0f894493-e622-4418-a1e9-097317145de2 · outbound

This paper cites Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models.

Accelerating Attention with Basis Decomposition Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.977946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.977946Z digest=sha256:5b38c1f74becbc3dd2c54fd38a7c897ca2f9848df461e9f9f04ad1df26a4eca2

Observation 6dd34c8b-9224-430b-aadc-c57a7b72dc78 · outbound

This paper cites InRank: Incremental Low-Rank Learning.

Accelerating Attention with Basis Decomposition InRank: Incremental Low-Rank Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.066165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.066165Z digest=sha256:5eeec963c9cfe71735583e2512bd98afc25aa1f7020437efe0c38e7406f20c74

Observation 2f8f7a76-881c-47a7-8042-894e98116cea · outbound

This paper cites write newline.

Accelerating Attention with Basis Decomposition write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.165106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.165106Z digest=sha256:e7533d3277897b20537506de16cda36af9fd42a8abd08383861c0d040febd66a

Observation 1d17f1a5-7d81-45eb-8f50-b2065d4371e8 · outbound

This paper cites @esa (Ref.

Accelerating Attention with Basis Decomposition @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.285194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.285194Z digest=sha256:8a49e4e94acf798935ca8a26add506cf9f97c8041e711717f56a95cd7b4f00a8

Observation 93e27143-3258-49e6-9bd8-d9078579c549 · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.403613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.403613Z digest=sha256:71a851b16e7f06d0026cd41fe7084997659106fc7541fbd7c307c0e3645b0160

Observation 07ee1195-c7ea-4742-86e3-af5f4901c69b · outbound

This paper cites u `:^!t )GeuwokcJ _ ]n?ICq .WT +BCBC &q=2.

Accelerating Attention with Basis Decomposition u `:^!t )GeuwokcJ _ ]n?ICq .WT +BCBC &q=2

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.489014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.489014Z digest=sha256:601a8f172c596a51dc024c26b1970af43956d3f15f28b3dd23a60ae5a90120c6

Pith citing papers

No inbound Pith citation observations are available.