Pith. sign in

Paper Citation Record · LEDGER

Accelerating Attention with Basis Decomposition

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2510.01718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01718 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:54:36.489014Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73142ff2-4f01-4393-9534-36de299848ad · outbound

This paper cites Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman.

Accelerating Attention with Basis Decomposition Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.229619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.229619Z digest=sha256:1ba142270f9b8d2f659a0dfb9951984323d70bb0761b939aa78c821d8c9f84ef

Observation 640eeb3a-151f-454f-9f2f-4da6f2fa88f4 · outbound

This paper cites Longformer: The Long-Document Transformer.

Accelerating Attention with Basis Decomposition Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.269838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.269838Z digest=sha256:f09ac4a899aad1fb01286c904929185a6efbe809768f46c547754bc7c961d518

Observation a7045ed7-be93-4994-92ba-59ddb6e45bc7 · outbound

This paper cites Language Models are Few-Shot Learners.

Accelerating Attention with Basis Decomposition Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.322033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.322033Z digest=sha256:978d52dc3dfd3a45e5d1017df75c01dbde04254bfa81e6da5e6f6de8c58bd77a

Observation 00d8f182-ab52-47d9-8058-734c395a446f · outbound

This paper cites Linear least squares solutions by householder transformations.

Accelerating Attention with Basis Decomposition Linear least squares solutions by householder transformations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.370688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.370688Z digest=sha256:6ccfa3e1dfaa3d58b451e729bc23dcde7934d8d9c5ed1352f379a6ba9e6cc3e2

Observation 5a81a44c-5501-4bb1-a19d-95d6a2fb30ee · outbound

This paper cites u ker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Marcello Federico, Sebastian St \.

Accelerating Attention with Basis Decomposition u ker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Marcello Federico, Sebastian St \

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.457634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.457634Z digest=sha256:625ca48fb3f8a8214b5ffc56155dc3cd5dc135a743952774d1fe620319577fca

Observation 5af29c4f-d56f-4c78-ac2c-3a8c2f8f77cb · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Accelerating Attention with Basis Decomposition Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.547368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.547368Z digest=sha256:49c0d5f3b0d5373f97c20fe45c4319b12d2088e00d4d3d2103433a1e2688a299

Observation cce563b0-900f-4561-a90d-b07ee1fdde09 · outbound

This paper cites Rethinking attention with performers.

Accelerating Attention with Basis Decomposition Rethinking attention with performers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.602091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.602091Z digest=sha256:828436ed348f99f9462517ccc6bf7c19ad30ceab8c87aa66a2beddf48557955a

Observation 20be47fb-d94d-4b85-8eb5-9b30cf963f76 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Accelerating Attention with Basis Decomposition FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.663535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.663535Z digest=sha256:a50a9edab884c3676757253f4e3a689e17e4373ec455da9d31f9f805392b484b

Observation 21a61534-ad57-4eeb-a424-c151aadd6afb · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Accelerating Attention with Basis Decomposition Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.706584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.706584Z digest=sha256:b2555015ac200bb8340d8f40fb822c53d7f548e87cb45d3b613868077a8ba81f

Observation b6941ade-a66e-46b3-b115-1816909e2f36 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Accelerating Attention with Basis Decomposition An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.763214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.763214Z digest=sha256:370952256eace6798d1f81338fadd0bf5ae005b3d3994733d5723203076d2f20

Observation a2bba1f1-1e01-4127-83a6-4f2bc1d452d3 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Accelerating Attention with Basis Decomposition Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.802792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.802792Z digest=sha256:361f1e36fbd87a662eac879b6fc794b5574b75a08091f7bfa366843c6a9f12a7

Observation 6010a440-46a5-4e74-89ff-412c8e4fa407 · outbound

This paper cites OPTQ : Accurate quantization for generative pre-trained transformers.

Accelerating Attention with Basis Decomposition OPTQ : Accurate quantization for generative pre-trained transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.824068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.824068Z digest=sha256:0e1e117324e1b2a73ebe56a02a6084934bfe752b5033cea202a7d94da25c1fa2

Observation 5a9db1f6-3352-4f02-9ef6-f5ca335451d4 · outbound

This paper cites Strategies for applying low rank decomposition to transformer-based models.

Accelerating Attention with Basis Decomposition Strategies for applying low rank decomposition to transformer-based models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.865017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.865017Z digest=sha256:3f293111a5bd605b4a2a1c2fa61cf98f5462f3551be94cd21a6aa258a3f967cf

Observation a057aba3-d32e-4e15-9bf0-b17222eaee1d · outbound

This paper cites SLTrain : a sparse plus low-rank approach for parameter and memory efficient pretraining.

Accelerating Attention with Basis Decomposition SLTrain : a sparse plus low-rank approach for parameter and memory efficient pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.910134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.910134Z digest=sha256:5c7b5f00578d0eb14842ec2f7f740688f6d983709b9a37d4c0578aebbec02257

Observation be4e18b6-074b-42af-9377-38971a8d2709 · outbound

This paper cites Language model compression with weighted low-rank factorization.

Accelerating Attention with Basis Decomposition Language model compression with weighted low-rank factorization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.931321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.931321Z digest=sha256:18c5fc5d10106254771b4c136c1f077666843fd16a9b057d946dc1981ff618e0

Observation 2a63d98f-6da0-464d-9196-37cc203100a1 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Accelerating Attention with Basis Decomposition Lo RA : Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.029756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.029756Z digest=sha256:554bb6df1def89aecef60971ee39cb8032b25ec56ebd5ff2f4e10c0fa24af3aa

Observation 2b676826-9b17-4b04-b090-538dc0d1f1d4 · outbound

This paper cites From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications.

Accelerating Attention with Basis Decomposition From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.091978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.091978Z digest=sha256:6811f4c99b8008dc44b8c128493598313b24d32ee66f7d9fddf379fd685ac4ba

Observation a6a14013-6408-4d60-858d-e52a7e75f9e7 · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

Accelerating Attention with Basis Decomposition Exploring Low Rank Training of Deep Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.209617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.209617Z digest=sha256:eb5b602136a82535f4bf309f9f3872e8dfa93058138f1708bf9f555aec6c4ee3

Observation 631783ef-d799-40a2-8984-3a39a3333220 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Accelerating Attention with Basis Decomposition Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.333202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.333202Z digest=sha256:c8bf3167adb69d6936ca726848adda660bee17532d829176580f5c3b2ba9d669

Observation dedb3b3c-0143-447c-86ac-5a807bdcfb2e · outbound

This paper cites LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression.

Accelerating Attention with Basis Decomposition LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.437126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.437126Z digest=sha256:f61cb69fe83ceed017bffa457a7411ada63e8a530d8dcfbe523eb89c25892285

Observation ec62ac58-2159-4e00-978d-e4bd905ef6e1 · outbound

This paper cites Tenenholtz, Lester Mackey, and Nicolo Fusi.

Accelerating Attention with Basis Decomposition Tenenholtz, Lester Mackey, and Nicolo Fusi

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.557629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.557629Z digest=sha256:2e965e2ffea1d30d65d015b2ea5b268995b5db30df203d94ba64c0ee9b5abd74

Observation 0f83de22-e96d-4476-9f62-7053a68070a3 · outbound

This paper cites Reformer: The efficient transformer.

Accelerating Attention with Basis Decomposition Reformer: The efficient transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.651506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.651506Z digest=sha256:42f9d174d781a1c6ba30136805e5c742de4266eb7fab525e640fa6f7e370a30f

Observation 31db1703-5b8b-4a93-b2ca-a68b152ecc44 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Accelerating Attention with Basis Decomposition Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.752334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.752334Z digest=sha256:34246236f007ca5b7c5255bbaa638da98d04d0b190a3918820e38e9a3afc43d5

Observation 643ba7d0-1136-467c-b008-08c2923f039d · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Accelerating Attention with Basis Decomposition Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.894154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.894154Z digest=sha256:7160e0e593f031f0e1abbe1f219bf3b4d9afc898eda3efee2afb0b92d561977e

Observation 01fa899c-3f53-4895-be6e-7047fb963ed7 · outbound

This paper cites L o S parse: Structured compression of large language models based on low-rank and sparse approximation.

Accelerating Attention with Basis Decomposition L o S parse: Structured compression of large language models based on low-rank and sparse approximation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.978012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.978012Z digest=sha256:98c24d0e512fe9abc13fbc3b22d33ed96a77a48ef8d94794855e0a1d29936d3d

Observation f261d3aa-14a0-4a8d-9c5a-d529ae8ba1dc · outbound

This paper cites Relo RA : High-rank training through low-rank updates.

Accelerating Attention with Basis Decomposition Relo RA : High-rank training through low-rank updates

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.133004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.133004Z digest=sha256:5fa91f4c916b36099d4aaba1aa77a6bab3cbe421f4ad8a98c793b19f4f4c4a2f

Observation 77cdf235-0bfd-4f03-99a5-b557d1825d12 · outbound

This paper cites MoDeGPT: Modular Decomposition for Large Language Model Compression.

Accelerating Attention with Basis Decomposition MoDeGPT: Modular Decomposition for Large Language Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.247901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.247901Z digest=sha256:33887f05ee8d507d26f85f6a5598fbe6307157371085f6e00b61e83330fe47f9

Observation 83137743-3cb1-4ff3-b11a-6cf1c74487d4 · outbound

This paper cites Duquant: Distributing outliers via dual transformation makes stronger quantized LLM s.

Accelerating Attention with Basis Decomposition Duquant: Distributing outliers via dual transformation makes stronger quantized LLM s

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.347427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.347427Z digest=sha256:dd7c1ea10985e45f0a9e5b1abe3ce88642a7e633418e1b0e1675bdcd7cbad3e7

Observation 9f0d11ae-4f6e-4531-a7cc-a1a6c23fc160 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Accelerating Attention with Basis Decomposition Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.471274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.471274Z digest=sha256:521d964db6f4f18233aad492a25c7f2c22d53e91226bc1e8d59fd6fe694afe15

Observation 5ae5674c-ffb0-4b20-839e-613f321f3d60 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Accelerating Attention with Basis Decomposition DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.527232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.527232Z digest=sha256:22f743987d7b12cce3e1213ff0df28b52e194be13bbaebaccf88da42349aef27

Observation 07b3b64d-3b08-49e8-8f1c-d43f6e3b80c1 · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating Attention with Basis Decomposition DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.615317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.615317Z digest=sha256:2a5169509644cf152de526175001da5535d9bbf4c7f7a96a0dbe66562c60734b

Observation 0258c38e-9abe-4a8a-b13e-c7ebc7dac914 · outbound

This paper cites Dora: weight-decomposed low-rank adaptation.

Accelerating Attention with Basis Decomposition Dora: weight-decomposed low-rank adaptation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.698827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.698827Z digest=sha256:628bd1e901bf42a72f8a0f3e3d96f363630a6511619aaff4c22c6f8f9b5b67ee

Observation dd5384e6-bb56-4bcc-8dd3-d47a66c0135b · outbound

This paper cites Eora: Fine-tuning-free compensation for compressed llm with eigenspace low-rank approximation, 2025.

Accelerating Attention with Basis Decomposition Eora: Fine-tuning-free compensation for compressed llm with eigenspace low-rank approximation, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.861188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.861188Z digest=sha256:959c01b604771edbd4b8e801ea2b666230c11c6b441ed74f7f7cde4838456623

Observation 7ac60136-ac83-4676-a145-dda4bcac9790 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

Accelerating Attention with Basis Decomposition Llm-pruner: On the structural pruning of large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.977760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.977760Z digest=sha256:b6bd6307bb8bcad4cf98584bafd1cda86de6a6fde0c4dc826170e20e33c83e8a

Observation 953b5e5b-d8ec-4d75-a517-f86fd2b5d80d · outbound

This paper cites Pi SSA : Principal singular values and singular vectors adaptation of large language models.

Accelerating Attention with Basis Decomposition Pi SSA : Principal singular values and singular vectors adaptation of large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.097917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.097917Z digest=sha256:238f904c54cbda4e3834f9fd9883ba5823be6898bfd5d0771548d082209fff98

Observation 3dbea6ef-0b63-4017-9997-9c4ea74232fe · outbound

This paper cites Accelerating Sparse Deep Neural Networks.

Accelerating Attention with Basis Decomposition Accelerating Sparse Deep Neural Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.214208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.214208Z digest=sha256:916829a54aeb876e4f5efc0890166e2f393d836cb9090ca97dfc2fe97754b451

Observation d439b2ad-6e08-4225-bfbd-55466bb6f4e4 · outbound

This paper cites Dobi-svd: Differentiable svd for llm compression and some new perspectives.

Accelerating Attention with Basis Decomposition Dobi-svd: Differentiable svd for llm compression and some new perspectives

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.322566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.322566Z digest=sha256:6f660395f3a07a6a6e6a2948b728efdad250bfd260e626df6e0116c81797e7bf

Observation 6f961544-b9b2-4dd2-8839-cc51cc6df62f · outbound

This paper cites Improving language understanding by generative pre-training.

Accelerating Attention with Basis Decomposition Improving language understanding by generative pre-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.440915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.440915Z digest=sha256:f2778e0133b46cf0a28e7ad8479d7449e592f631545c3c07f1a43fc0aacc40aa

Observation c43b9212-5043-4bf1-9d40-21ba7e391419 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Accelerating Attention with Basis Decomposition Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.536976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.536976Z digest=sha256:b9579230aa10f0a09fdf027d33cc9c3da65e9c6ee0d10edc05cc63c301c993ea

Observation 278dc369-fde2-4308-9668-e49d0114f923 · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.587104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.587104Z digest=sha256:4257ad39d23708bee214fd8ac73f1e538f248286197c97d0eb257689099e2fd0

Observation 31e01949-78d7-4159-beb9-0126df88ca27 · outbound

This paper cites Compressing large language models using low rank and low precision decomposition.

Accelerating Attention with Basis Decomposition Compressing large language models using low rank and low precision decomposition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.690785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.690785Z digest=sha256:82db614c48f36894ce396749eb2db5282d27ec8780787c7e578acd1c53912b1c

Observation 342da8c0-a4e9-4b3f-9574-f2d413e0f6b4 · outbound

This paper cites ESPACE : Dimensionality reduction of activations for model compression.

Accelerating Attention with Basis Decomposition ESPACE : Dimensionality reduction of activations for model compression

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.809317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.809317Z digest=sha256:21cce41f6a1bb097694d62ed8585e7b529db399b59c6255a579ae34627ab3555

Observation e920ecbc-9788-49b6-8e95-0174b00261b3 · outbound

This paper cites Robust low-rank training via approximate orthonormal constraints.

Accelerating Attention with Basis Decomposition Robust low-rank training via approximate orthonormal constraints

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.919849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.919849Z digest=sha256:c468c85ad1f3c2e85c3d9f68f64a16f34a406c14d2c30f5005dcd606c6ccf574

Observation 1ea1abe1-5963-4ead-a7f4-2d53663b6ac1 · outbound

This paper cites Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations.

Accelerating Attention with Basis Decomposition Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.055161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.055161Z digest=sha256:0a5aad53ce73fb0ca447ea0dfdf8fcccb469bbaee246604ef0697ba790f27c24

Observation 93e826c3-8a4c-4eb0-a545-a663f681d28b · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

Accelerating Attention with Basis Decomposition Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.234225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.234225Z digest=sha256:7d612e70031dd89df4d36d0a98da2658801a3b32d791010e28831ff9b3e2cfc3

Observation 586d5a36-5a80-422b-b19e-72f3bccac425 · outbound

This paper cites The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction.

Accelerating Attention with Basis Decomposition The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.403786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.403786Z digest=sha256:d540cbb8e7df644235947051bcc13273fd638fb16a44dbac95cb8e0fced27229

Observation 528f9003-faa7-48b4-80d5-90af2d766239 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2021.

Accelerating Attention with Basis Decomposition Roformer: Enhanced transformer with rotary position embedding, 2021

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.560871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.560871Z digest=sha256:73d9e29cc256bf83f8c1f4889cd5f825b08dceeeb6a678d507729e1d46c72c66

Observation 18fd2129-8fbe-4609-949c-f15b7ecfa635 · outbound

This paper cites A simple and effective pruning approach for large language models.

Accelerating Attention with Basis Decomposition A simple and effective pruning approach for large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.725960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.725960Z digest=sha256:c19417bc9a7773994cba16b4515d22acc25589680c53cdb2d4ec7f31d6c2a212

Observation 3ae9a8c8-8529-4874-9453-f8d373e940e4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating Attention with Basis Decomposition LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.850727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.850727Z digest=sha256:4d07060117ab9837f0bb83474667de064e060da10d8a9e249d3d5b7904e5f7bd

Observation 3e2ec9d2-3863-4724-a5b5-713fdfd1aa2f · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.978563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.978563Z digest=sha256:33cfed1627738ea44cd9dce40f102264b5e9e96f0d7313b662ec4bb686dfc8fa

Observation f0136548-5b07-446f-8cd5-045e713c7f8b · outbound

This paper cites Attention is all you need.

Accelerating Attention with Basis Decomposition Attention is all you need

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.063112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.063112Z digest=sha256:a60078dbbb4748ba15cd913363bd752ec67dbf1bd5167d1f1fe5f36fa50fd6fe

Observation bec3bf8d-1d0f-4873-9ac4-a9fea3897be1 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Accelerating Attention with Basis Decomposition Linformer: Self-Attention with Linear Complexity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.148840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.148840Z digest=sha256:fe57318aeb1527191631fb105a2deac3616697e2934cabe1ca1ef2729093a141

Observation d362884c-1a04-4ac8-93cd-c40cc7ca3094 · outbound

This paper cites SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression.

Accelerating Attention with Basis Decomposition SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.258159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.258159Z digest=sha256:09bec6d8e08270ecd9011fee8f68ade31dc56cba44bbd740952185f2e75eb869

Observation c321b518-a77c-447f-97dd-88711359f07e · outbound

This paper cites S mooth Q uant: Accurate and efficient post-training quantization for large language models.

Accelerating Attention with Basis Decomposition S mooth Q uant: Accurate and efficient post-training quantization for large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.361956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.361956Z digest=sha256:47339eb92e27a42cf6f29774d1fc3a5d88c5299e8c56072dc6d8c3f5b5c6be4b

Observation 3a849161-2923-4d55-bd1f-e3a7efe8544a · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

Accelerating Attention with Basis Decomposition ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.464572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.464572Z digest=sha256:a5daa9e01209a8b941bd3ed20447f0b4fc6dcec5f34e9abd5d8ae9eac8f63198

Observation fd338049-9301-4ef7-b7ea-92b6c9f09adf · outbound

This paper cites IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning.

Accelerating Attention with Basis Decomposition IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.581040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.581040Z digest=sha256:d45a3e513a891f03bb0f75e8f7a51306526073982dfc6060687219978f664210

Observation a45d9535-e250-4bd4-bb31-6d366b7a860a · outbound

This paper cites Adaptive budget allocation for parameter-efficient fine-tuning.

Accelerating Attention with Basis Decomposition Adaptive budget allocation for parameter-efficient fine-tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.660902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.660902Z digest=sha256:fe99197f1a22e2a476eb34311d221c55dc1c2ea2c83e0eb262bd55b2f92dc568

Observation b1bce96b-694c-43aa-9b5e-4f059156fc26 · outbound

This paper cites OATS : Outlier-aware pruning through sparse and low rank decomposition.

Accelerating Attention with Basis Decomposition OATS : Outlier-aware pruning through sparse and low rank decomposition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.775920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.775920Z digest=sha256:65b2a15c4390d6970a4164920fd1d639309eb5012b271dc5466e0e7f403f994d

Observation dc3d7516-ad5f-4dd5-a41c-de003089dfb9 · outbound

This paper cites Plug-and-play: An efficient post-training pruning method for large language models.

Accelerating Attention with Basis Decomposition Plug-and-play: An efficient post-training pruning method for large language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.900097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.900097Z digest=sha256:07c2ae8f0348fe5a59d2e4f7b1fb30bc2bc7e5ea693909d52e526cc28980f467

Observation 0f894493-e622-4418-a1e9-097317145de2 · outbound

This paper cites Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models.

Accelerating Attention with Basis Decomposition Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.977946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.977946Z digest=sha256:8e4c9ca4cfb06a72f56c78075a008c683cc7a5e8848af7addb60710e944238b6

Observation 6dd34c8b-9224-430b-aadc-c57a7b72dc78 · outbound

This paper cites InRank: Incremental Low-Rank Learning.

Accelerating Attention with Basis Decomposition InRank: Incremental Low-Rank Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.066165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.066165Z digest=sha256:db56d2cab81138dbeb4be7e4ba4bd98ba7256d2770f034a46fa8f4f7a590d433

Observation 2f8f7a76-881c-47a7-8042-894e98116cea · outbound

This paper cites write newline.

Accelerating Attention with Basis Decomposition write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.165106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.165106Z digest=sha256:f6d9800f13171779318d7d15a3c9f6647b7978b2ea6cac735c7ac340f9ca52f4

Observation 1d17f1a5-7d81-45eb-8f50-b2065d4371e8 · outbound

This paper cites @esa (Ref.

Accelerating Attention with Basis Decomposition @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.285194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.285194Z digest=sha256:650986b98137c9098b8101b019e59d22f548039648ecaf7c4ce1a84daba15744

Observation 93e27143-3258-49e6-9bd8-d9078579c549 · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.403613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.403613Z digest=sha256:e24d99385309a4c0f6d1e9f3bdac1491b32eba236a797354c2858c18711ba806

Observation 07ee1195-c7ea-4742-86e3-af5f4901c69b · outbound

This paper cites u `:^!t )GeuwokcJ _ ]n?ICq .WT +BCBC &q=2.

Accelerating Attention with Basis Decomposition u `:^!t )GeuwokcJ _ ]n?ICq .WT +BCBC &q=2

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.489014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.489014Z digest=sha256:7e9a36ad1b1436751708152fb4c0eead228f2c8ba8570c811575e13dfb0dc7bf

Pith citing papers

No inbound Pith citation observations are available.