Pith. sign in

Paper Citation Record · LEDGER

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention

As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2506.09316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09316 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:56.185517Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T01:44:40.722957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T01:45:50.813321Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bfda6f3-306d-41fd-931b-8187ceac2d05 · outbound

This paper cites write newline.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:47.722470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:47.722470Z digest=sha256:26de0736e2079cd8e6bab5780ad88bee7f952db50967287fc96d3917852c15fc

Observation f13a33ab-b5eb-4658-b2f7-6f3d8873d112 · outbound

This paper cites https://huggingface.co/blog/bamba.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention https://huggingface.co/blog/bamba

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:00.985710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:47.896982Z digest=sha256:18d0207b49ce921af2f51d203b338c40839c5dffdc0510fc8e642c4201ad32f3

Observation 0b7d6472-f4b2-4a0e-a991-5a5a52832d11 · outbound

This paper cites nvidia.com.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention nvidia.com

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:00.736487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:48.049768Z digest=sha256:352b064387f206c7a6741bca1325dc6aaeed82e951cbb82e861a86a9ca50bc0b

Observation b283f60b-528c-487d-937c-6b79e056f232 · outbound

This paper cites Y., Rajbhandari, S., Awan, A.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Y., Rajbhandari, S., Awan, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:00.471006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:48.164225Z digest=sha256:249d25460bb0c26e07fc40336abce6d2f57a002012824ef9594759d520dd848a

Observation 98ce186c-6273-4abe-aaad-b75d1ef8b7a1 · outbound

This paper cites E., and Pedram, M.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention E., and Pedram, M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:00.081477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:48.249807Z digest=sha256:a7f6a4e5fc72b747c59729bef405e142f47f2e9122ad936d02e93d17e8e10e5e

Observation 5273c4d9-aed8-4fc5-9d84-7a4e9a747c66 · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.382295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.382295Z digest=sha256:b9be1ab47aa748f952741a2aa88a936ac729f3649d552a516b5d3c8e69610517

Observation b4b23c6f-a9c4-450c-af41-b565de419c9f · outbound

This paper cites Longformer: The Long-Document Transformer.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.482440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.482440Z digest=sha256:a38409602dc420dfc1bf99ed99ec8691c42e1a2daeed0898f60ec64187ecd34a

Observation 3b11432c-7289-4214-a0e8-0356dc373cd5 · outbound

This paper cites DeciMamba: Exploring the Length Extrapolation Potential of Mamba.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention DeciMamba: Exploring the Length Extrapolation Potential of Mamba

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.617700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.617700Z digest=sha256:bcdc0097264415d136818f459714326e07e49b00d62e3f8656625c00b539e226

Observation eaa20003-4262-4252-8800-37325f79b0d9 · outbound

This paper cites Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.707551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.707551Z digest=sha256:a33e1789ce7271adebccbff1bde3c8f421ce053aeb0b8fb981de68138e3d1298

Observation c038ba8e-116a-48c4-9d24-081a163a488a · outbound

This paper cites Read-me: Refactorizing llms as router-decoupled mixture of experts with system co-design.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Read-me: Refactorizing llms as router-decoupled mixture of experts with system co-design

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:59.780840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:48.781558Z digest=sha256:19afb33c22385c56da2245e9bbaa2c0e1d9f27239ad7d2543fd1b0a0ae26c238

Observation 07d2a5a3-58fd-4863-bd67-88bfea0b78d1 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.869443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.869443Z digest=sha256:b59b9c791cab7fc5cf7905da83fe6af9be6460eb11e0e68882698a375c705169

Observation 1c3c1ef5-6963-4f2d-be24-b2e6f1caf26a · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:48.971053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:48.971053Z digest=sha256:801df00e84683e6ef09d88685e18ca150dd4e410fc790ad8b769dea6960d2196

Observation 3f981207-6675-465c-bddf-5865bbf9af51 · outbound

This paper cites Rethinking Attention with Performers.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Rethinking Attention with Performers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:49.052883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:49.052883Z digest=sha256:b048236e408f16aaaa43c3506c184ebc5d66d06093394e9fc147d4c9c21d853d

Observation caa8a0d6-4443-4244-be32-5541b466d2b0 · outbound

This paper cites B., Bierbaum, M., O'Keeffe, K.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention B., Bierbaum, M., O'Keeffe, K

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:59.552638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:49.197097Z digest=sha256:7d6b563692237f4019cc3f9f218a2c26f4622a00178b50b887d30f281ce1a202

Observation dd369b7c-a24e-4e77-9095-027a32d5b3af · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:49.317548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:49.317548Z digest=sha256:9146a7b486be5881ea133688926b550befe4a92b791e8ca349e1b57fbdc4f6e2

Observation 7656668f-69cb-4452-a0d4-fd5e041f863f · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:49.481739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:49.481739Z digest=sha256:c004cfd7dded41e28be40af1a2c07f86fe8093974519af80e97c17265f0f737b

Observation 55ab2dfb-5b33-4648-897b-1a4f4fa7ad5f · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:49.567474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:49.567474Z digest=sha256:6afe8d660fe51902176afa31e7a5f993fcd9b8cd9686d4cdd4af4057d2d169b6

Observation 131ca01e-8ce2-4a38-a120-994baca3c13d · outbound

This paper cites Stack exchange data dump, 2024.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Stack exchange data dump, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:59.265908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:49.666529Z digest=sha256:2f83ae532a0bcfe22625cb1e9ee6f055b11f6cb9ca22b8fdf34eb36b85f5f50d

Observation 64f49872-e4ba-4ea9-81e7-b1f0040dceb4 · outbound

This paper cites Wikimedia downloads.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Wikimedia downloads

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:58.999314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:49.775365Z digest=sha256:2d436fa04bd0dbf6418b7d5dea17ec0a05411d7aabcecbdd08790dc154bfc11f

Observation 4e870e7c-e10d-49bc-a9a6-3d37f9ae665a · outbound

This paper cites and Alistarh, D.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention and Alistarh, D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:58.788465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:49.889850Z digest=sha256:e789a068ddcd5ab6a497ed876fe09fa0c15442d09522ca9c5b98897d7b65466e

Observation 9b73c7ec-2a49-4a10-abcc-e60a0b7ec9ee · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:50.021067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:50.021067Z digest=sha256:feac11ac8a4c12b82725ec25a7e2c12d1ce79a97a0722f39f7aa1dd2c1a0423b

Observation f2f1a551-1433-40be-83a9-31b56dcc05cc · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:50.184312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:50.184312Z digest=sha256:f1d3ba946ed06f770880ce1431f20c2fbf1c8217bfa6c8e3f11a51f12e54e771

Observation 00105b00-2053-4aee-8bf6-8b37ef7c0782 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Zamba: A Compact 7B SSM Hybrid Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:50.356463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:50.356463Z digest=sha256:c6ce09a01dc1d32a3a582785f79849e76ce2c68f2b3826de4bbc5c345ddfaab0

Observation b572fc40-30bf-400c-89ce-597be1886fbf · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:50.507722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:50.507722Z digest=sha256:25db2a7b62e6cdb32359e1e981ed9bff1ee2788565a9249768a72384877e83b9

Observation 4df2d4b4-e48d-4630-a050-d2a0ad4d3c48 · outbound

This paper cites N V I D I A H 100 T ensor C ore G P U D atasheet --- resources.nvidia.com.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention N V I D I A H 100 T ensor C ore G P U D atasheet --- resources.nvidia.com

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:58.550360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:50.652089Z digest=sha256:968a1b93d5c7870f1d1a5d0d5a9594d5d9dba380b254fbd163d0e6722f516db0

Observation 24fec936-f6af-457d-afa1-96c547d714c2 · outbound

This paper cites M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:58.326820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:50.783539Z digest=sha256:a3edc0260059a26e3a5c67b8ea4f439504fd0f9bf8e8bafd93687bad45bf4d7d

Observation 2af91c23-f9c4-4c00-adfb-ca5d731a8b56 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Distilling the Knowledge in a Neural Network

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:50.924458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:50.924458Z digest=sha256:d26a3dd6458b4385be9e2a05f05000b162b7d8a4d92e1b91cbaca52cc506e1a9

Observation 474fa1c2-123c-40e7-83e2-405f0ef97487 · outbound

This paper cites What does bert learn about the structure of language? In ACL 2019-57th Annual Meeting of the Association for Computational Linguistics, 2019.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention What does bert learn about the structure of language? In ACL 2019-57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:58.092789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:51.094449Z digest=sha256:c83ac48451a398e7cb75c9bcb9d86cb6c26272bcb5c06416ef2fd8ac513519fe

Observation 9c07b291-0a01-4b32-97a8-91d25c9f1e2f · outbound

This paper cites Finetuning Pretrained Transformers into RNNs.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Finetuning Pretrained Transformers into RNNs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:51.297853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:51.297853Z digest=sha256:bdc7e62e590eae19f6032b558758a5085ff6386ee429d9ef26e5de327f224e6d

Observation f4fd3083-54f4-497e-9cb9-71726f172800 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:51.450640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:51.450640Z digest=sha256:2690f1d91bea4167cb43a411890ded1fcd1ab646a5243e59eec17b10722b2766

Observation b7bd5cda-0fe8-4210-9584-91b6f27db0bc · outbound

This paper cites MatFormer: Nested Transformer for Elastic Inference.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention MatFormer: Nested Transformer for Elastic Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:51.628661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:51.628661Z digest=sha256:511fd7a21cb14c9aa2ba7f84bf7c70b68f3bd24c0f2dc040186bd0e14309497b

Observation 486207a5-8545-41c4-a53b-f6594aeacc91 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:57.806373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:51.721497Z digest=sha256:af1ba0e1ef83d7194a273b445cbaed3cd9e163e45e7be5564d06c4d8d1c668f2

Observation f8dbd3fc-40c9-40ed-9b15-0bc4e9df16b3 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:51.906431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:51.906431Z digest=sha256:4a89489d2e5356af109f90be3c6151bc541bc3205abf5fec161553991a4c4f89

Observation b52c1bab-5b63-4813-a562-1d957af6f787 · outbound

This paper cites Fast inference from transformers via speculative decoding.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Fast inference from transformers via speculative decoding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.096427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.096427Z digest=sha256:7bde6dcb7d686008634438f2cbb969bdc64ab3ce11605acce995a02f333d4065

Observation 2a74723d-51c5-4d7a-823a-b6da73242916 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Textbooks Are All You Need II: phi-1.5 technical report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.271444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.271444Z digest=sha256:cdd22a634bac01fb3a9f50afe237936fef205a14852d52b8f761192462541a09

Observation bf0c30b2-e33f-4a74-b80c-251072379c79 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention SnapKV: LLM Knows What You are Looking for Before Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.402890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.402890Z digest=sha256:36148093c2c3968e048e9b8f3247775ec17eca0f61ce6befad7ee2f92919091b

Observation ee832730-6ad7-4ea4-9774-9acec88298a3 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Jamba: A Hybrid Transformer-Mamba Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.520878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.520878Z digest=sha256:c4a5c149a0aecf6e606712d158d9eb302fd40900c726846e70697e264e4341b2

Observation 0cd404f9-d719-4b20-970b-990fd9feea19 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Rouge: A package for automatic evaluation of summaries

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.577018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.577018Z digest=sha256:72cd3d530e01cd5a6be2ad1d875c2699d10a16253d13b81cab34324cc6a6ab79

Observation 9eeb5ed2-3dd4-41f4-a0bd-08b52facb7b3 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.659154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.659154Z digest=sha256:669cc3deaf85bd617fe8aa7e26df9df0718b15af7663a9e0d991b3cc4106bfae

Observation 02c42fcb-6b4e-4151-b79c-b91655505b04 · outbound

This paper cites Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.744221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.744221Z digest=sha256:fbfad0387f39702d47e2e9affd56e77eabcd16d97345c31bf0dd843a93e5e7b7

Observation 21557610-a2cf-4868-8bde-b8463e0986e7 · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:52.920517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:52.920517Z digest=sha256:cdcf9512829dcbabc6e6197234411a37d9826f54fbdedb6ebf07474c8855f6fd

Observation 8b5a5c8a-968c-454b-8ef8-e040a5dc07d0 · outbound

This paper cites Decoupled Weight Decay Regularization.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Decoupled Weight Decay Regularization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.029160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.029160Z digest=sha256:87192cf74d37dfe4b64028c9d2e630e1fc9bdb62a99e9d0a6609247d8a853140

Observation 22cd5a87-b673-442a-8b2d-2bf833df611b · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Llm-pruner: On the structural pruning of large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.116479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.116479Z digest=sha256:37db232820569c448b8ff9dc7c2ff9acc66ecd888d732ce8acc4dee6a4f5eafe

Observation 0b665a0f-8f6e-4ac9-9596-5380efcd79bd · outbound

This paper cites Linearizing Large Language Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Linearizing Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.185985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.185985Z digest=sha256:77fe7d5fe2431aa15954e9980eba97d8f14f8619449b443965daf9c063a4c591

Observation 12c7d4d8-c4d6-4ec6-afda-aebf182d97fa · outbound

This paper cites Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.298370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.298370Z digest=sha256:9ce7dbcca6f88bb7d512ba1d0aeff2ea590fcf7b01d511a732af9cc8f6948a60

Observation f818956f-faa5-4580-8650-0a35e4044183 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Splitwise: Efficient generative llm inference using phase splitting

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:57.439523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:53.393094Z digest=sha256:831fe64f8018f76967b58d096bd228bfbf210b6a3dc4a3ee0aa9797041e51f3a

Observation 1a8c189a-0b4c-4e2b-a3ff-5b961b0940fc · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention RWKV: Reinventing RNNs for the Transformer Era

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.446107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.446107Z digest=sha256:d58c59f66d197b06d7b7030da371764140bd0f8be0311364aa9cd957e5e0f508

Observation f69b21ac-6d7e-4951-9be5-a155f70a92a3 · outbound

This paper cites an unresolved cited work.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.510273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.510273Z digest=sha256:5ecadc9d9ee799611179a275e7779bb6f1a37c851128f1f1104988d5bd6272ac

Observation c20b2571-24fe-4f0a-8c43-4f5d0a729026 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.560687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.560687Z digest=sha256:d3fdafebe6bb319b5ec1a15397379e9e911feed58658ac46bc2871f8e9439171

Observation 5667ba96-4ed2-4d78-a3d6-b58cc4f15d35 · outbound

This paper cites Optimizing transformer inference with selective distillation: Layerwise conversion to linear attention.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Optimizing transformer inference with selective distillation: Layerwise conversion to linear attention

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:57.163885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:53.630222Z digest=sha256:d91e041264926baedbc07652b4a69a12e314c08f0cf85bb1c3ff6009afc4ef15

Observation 7438f84c-e48b-4583-94ed-4f2777083c0c · outbound

This paper cites Get To The Point: Summarization with Pointer-Generator Networks.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Get To The Point: Summarization with Pointer-Generator Networks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.681251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.681251Z digest=sha256:98d99a4fc4a3ecf63685638c48fa8ac14da12b17707b20e3b44343e58b7b6d27

Observation 19a7c27f-c1ea-422c-aba3-76d047304e95 · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.782070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.782070Z digest=sha256:d4ebe9b0341a594e9870cf899ad236ae693c4aa04f209856dfa10619bb6d6461

Observation db541a76-d744-4267-bcb4-3b42e9e53540 · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.855974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.855974Z digest=sha256:f9f515f5f9941ef1b8956088c3b48bb98552078f39951db609f85b6ff892aa8b

Observation c80eeb53-48e5-42f8-98f1-e8219f002748 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention A Simple and Effective Pruning Approach for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.946901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.946901Z digest=sha256:d815153a076e0d9c573b4c7253d2ef8baf1bd766260896f052426438af900bd0

Observation 4f864534-554c-49e0-b005-f2559d13cd72 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Retentive Network: A Successor to Transformer for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.032690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.032690Z digest=sha256:a6f4d7405ae8caae3cc6969a2e08a64905742d7e876616cf647a91026ff3dced

Observation 5ebb6325-398b-4bcd-8eb1-b762d92e1674 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.106435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.106435Z digest=sha256:b8b43c9241ea3ba7d87cceb6cb5b64bc07e02807dadcdb1a019d760fee6d9d5f

Observation cb530a10-3b22-4f0c-9a4e-a4d1d398d34d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention LLaMA: Open and Efficient Foundation Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.198656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.198656Z digest=sha256:cd767bcc6a3189da41b9ab5504e94c2c1df2c2f7ec07298d068e388017884584

Observation 9eae3fa5-0e4f-4652-baa6-55672d7c2c02 · outbound

This paper cites Attention is all you need.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Attention is all you need

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.287461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.287461Z digest=sha256:2e9c6ad3606c124841737e8bbabd70b1b54fcc43034d5e34ebf14bf20327202d

Observation 4bd7ea9a-c26c-40db-974c-73cf8af22a5c · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention An Empirical Study of Mamba-based Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.362778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.362778Z digest=sha256:ed497782e59c0ca31a656e0ecaecc49a5e61084024472588a2aa074c115c55be

Observation fd2e9bb5-50df-4e33-b77f-02ae198cc928 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.442743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.442743Z digest=sha256:061e906b12b2c89e40433388e7240e408d560631e3ce93b19e9fc4b5707bfdbc

Observation c37d3d81-be26-491d-8cf4-c007a1195ab3 · outbound

This paper cites The Mamba in the Llama: Distilling and Accelerating Hybrid Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.509957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.509957Z digest=sha256:e5fdca38f9832f96e800d4cad1fdf78ae94f2e99f3b7af2e01b0ad667182e7c7

Observation 78136a9d-755a-4e54-a2d9-3a851978ff3c · outbound

This paper cites Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.578841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.578841Z digest=sha256:7cc060ac71d988b9e54fdd14afc23c4b94cea241951685465758e6f74935e93e

Observation ac56f147-bab8-4dec-8882-d00e29c9e3ba · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Linformer: Self-Attention with Linear Complexity

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.669694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.669694Z digest=sha256:f91ba9322b71f18bdf2c8fdc63c725313c466b3bddbd63daadfcda47abb7ea32

Observation bd53dc6b-d00c-4058-bedf-aa578be4126f · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.748448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.748448Z digest=sha256:fc6982397a7c287b31e5f6ca957e0a54ee03fe358eac9abd39474d8d19ae6f91

Observation f53d8580-e9a3-47d3-8a15-896ccef73530 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Efficient Streaming Language Models with Attention Sinks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.815383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.815383Z digest=sha256:ac27a3c0063127c2893e648234675753e7249464f0de2b98386a096182f86018

Observation 508c89cd-b939-4ffb-8522-a5a5eb53a68b · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.918582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.918582Z digest=sha256:8fdf7a8bf0356641982f46266319f6adb7d7838aa2cd9c9e5fc87558428e107a

Observation f5a99ef4-a68c-4e5d-a02a-a37257d44035 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:55.007481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:55.007481Z digest=sha256:46daefa267ee89be62c65d2691b3b9a289a3d0b6c3d3c07058afa049472aba1a

Observation 806fe173-a090-432d-bc7b-b63718448c40 · outbound

This paper cites S., Kim, G.-W., Kim, S., and Chun, B.-G.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention S., Kim, G.-W., Kim, S., and Chun, B.-G

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:58:56.922471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:58:55.141581Z digest=sha256:93c939bc283b52daa038c74a051326f5370467069d92806a63624de7528b104a

Observation 9ffaab65-25e7-401b-b739-253b974d5ed5 · outbound

This paper cites A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:55.348157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:55.348157Z digest=sha256:872dedb5a67a84d42682bf07814c2206ef8a452d1168985bbf3e8fc40d3e2ca5

Observation 5c2ba22a-125f-4c20-b045-c6e43cb22b8f · outbound

This paper cites LoLCATs: On Low-Rank Linearizing of Large Language Models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention LoLCATs: On Low-Rank Linearizing of Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:55.594059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:55.594059Z digest=sha256:ab91aade490a2800fa54be29cf256d5ee46d8d6f76148cc46d91f5542b844189

Observation 09d1b7ac-3019-4ad9-81a3-0dd013c889d8 · outbound

This paper cites The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:55.827537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:55.827537Z digest=sha256:f2bd2139e778b36b4568a4b2f58c5e8853c6913b2760b95765aa1cc45256ebc3

Observation 397da5de-50cc-4ae0-8606-ec5277bca783 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:56.004667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:56.004667Z digest=sha256:0f501ca96eb2659f5ed9a19a4758981bfb35e0a56bd20ffb000abff923aca361

Observation 05c3c4a2-b8a6-463f-aa81-d4c5e9be40ae · outbound

This paper cites P., Zhang, H., Gonzalez, J.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention P., Zhang, H., Gonzalez, J

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:56.185517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:56.185517Z digest=sha256:aa8dc6c219c9da4fd9d3ae4f4cfa25b539797d1fc1ac3688ad1d14537f7fb75e

Pith citing papers

Observation a94b9bcd-0ab4-4d30-aa30-fd841611e3de · inbound

The Key to Going Linear: Analysis-Driven Transformer Linearization cites this paper.

The Key to Going Linear: Analysis-Driven Transformer Linearization On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T01:45:50.814563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T01:44:40.722957Z digest=sha256:89cbebf17df0d2643b75c9f8b3548a35ef7453caf4af6759482d350d866d249f