Pith. sign in

Paper Citation Record · LEDGER

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling

As of 20 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.24179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24179 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:10.193536Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1181232-b07a-4c73-a2fe-3033e696e651 · outbound

This paper cites Booksum: A collection of datasets for long-form narrative summarization,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Booksum: A collection of datasets for long-form narrative summarization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:27.136294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:01.544466Z digest=sha256:0af76dbfd0ddd526cc6321c946d309cd646bf515a6cfb6ce3041373501fdd291

Observation b13ab18e-98c6-42c3-beed-4206e077c3af · outbound

This paper cites Transformer Based Implementation for Automatic Book Summarization.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformer Based Implementation for Automatic Book Summarization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:39:12.172785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:01.698014Z digest=sha256:652bd13b3ab5c567b284279540aaed8b62d47c58244b94f2e3060cd66cc34cf2

Observation 512c5be2-fb3e-47e2-8561-02f765bf403d · outbound

This paper cites Booookscore: A systematic exploration of book-length summarization in the era of LLMs,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Booookscore: A systematic exploration of book-length summarization in the era of LLMs,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.814959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:01.862280Z digest=sha256:ce1e98ae9086a50d26e4947150f04700e2b50e57f69ea91503e93646c19748a4

Observation 8d871b43-4e3d-458f-adad-b036b31dd6f2 · outbound

This paper cites Peek across: Improving multi-document modeling via cross-document question-answering,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Peek across: Improving multi-document modeling via cross-document question-answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.607039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:01.995486Z digest=sha256:e02f4e5e22ee203ccb46251030ab79c22903f3b7f053ab271f96acd59b16a4b7

Observation 0492a883-5869-4d00-b33a-23eaf3cf4da0 · outbound

This paper cites Quality: Question answering with long input texts, yes!,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Quality: Question answering with long input texts, yes!,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.360575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.119346Z digest=sha256:4311c0d5b3d2ab600aa0dfd812e37a292568b0a3975136a17defec224f7d3fc2

Observation 51bd3625-2748-411d-8673-c8a17d439cb9 · outbound

This paper cites Eli5: Long form question answering,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Eli5: Long form question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.148743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.225120Z digest=sha256:dd069f6e226756d5770107d7746a93a25e9805bc211c27006e09c6851ded932b

Observation 11e139f6-855e-441e-9eb1-dee98a1dcd55 · outbound

This paper cites Teaching code llms to use autocompletion tools in repository-level code generation,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Teaching code llms to use autocompletion tools in repository-level code generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:25.930551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.354328Z digest=sha256:fd6eed42ba92debe331ced84a2f58aa36fc18c0e3d76a823bcaee6860d90fc3e

Observation d394e65c-6b84-4b07-85e6-84e86266e87e · outbound

This paper cites Rlcoder: Reinforcement learning for repository-level code completion,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Rlcoder: Reinforcement learning for repository-level code completion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:23.459473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.439280Z digest=sha256:faa15274f6cf15c4a5acc6e4c1e90a9a9fad7d53d10e7526795bcc8361b297b2

Observation cd7e2f28-2281-477c-ae9e-67a4033939f2 · outbound

This paper cites The llama 3 herd of models,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling The llama 3 herd of models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:21.967928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.540264Z digest=sha256:4d3efb289ffb0d7c1739947a8e248227a0b2d502fadb491361552eacab4d43ba

Observation 3dbd470f-d605-4360-a5f6-439816f638b5 · outbound

This paper cites Qwen2.5-1M Technical Report.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Qwen2.5-1M Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:02.628737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:02.628737Z digest=sha256:d51d1372962f0ea927b15202cbd3836a07024b7279d28fd51c09eb440d887e1e

Observation 7024eba2-1cc9-4e41-8937-4623bc89b760 · outbound

This paper cites Gemma 3 technical report,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Gemma 3 technical report,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.840840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.777619Z digest=sha256:f4e726904a551ec4f581ae886653ef0ab7b5cb6bf3a9006041d225e843298e61

Observation e6c144f1-0568-490a-8e08-926250a7c254 · outbound

This paper cites Deepseek-v3 technical report,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Deepseek-v3 technical report,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.572954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:02.883538Z digest=sha256:09167d9dc4785db205ae7d1ad651a51631a39e6de3667e16e355690417ac4d63

Observation 3ebd4d0b-4c23-4611-a066-be719b0961b9 · outbound

This paper cites Attention is all you need,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Attention is all you need,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:02.961058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:02.961058Z digest=sha256:91ba03a94345c38e3d5bb1b6b7b9c0f5e031da2157106912fd57d56b66c13fa2

Observation e8d31f66-3863-43e5-8cba-31d97e3a4ff1 · outbound

This paper cites Challenges in deploying long-context transformers: A theoretical peak performance analysis,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Challenges in deploying long-context transformers: A theoretical peak performance analysis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.317295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:03.066855Z digest=sha256:4e6bae348316bc1577fe241f12273a7e4418ef168565a9f6ace1e862f3c9f5dd

Observation dc30b61e-5109-45e4-afa2-82763aa26a9e · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.087881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:03.195707Z digest=sha256:6dfeb81187bb23012c52dae0f313f41b64c6879003aab2f91383d92c7314c61d

Observation 9b8de900-60eb-4bab-8ac4-71d127049db3 · outbound

This paper cites How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.324221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.324221Z digest=sha256:c25633f304beed5b013cdee4df74b793573012a9ee8d36dc68772476c23d3256

Observation f1fbed45-88d9-46ac-9bde-e090760fa7d3 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Generating Long Sequences with Sparse Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.429468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.429468Z digest=sha256:ce0c3a072f691a164e391382d4b97bc3a24ca3f82bda9d996f71c3643ca7ff63

Observation 3fb3dcee-4063-43f4-9f7a-9f4a119ed900 · outbound

This paper cites Big bird: Transformers for longer sequences,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Big bird: Transformers for longer sequences,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:17.709833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:03.569330Z digest=sha256:491cd96154c28e46dadc6959cc44dd24024ac3064ddb1404fb9c51ad64dd9652

Observation 1077f3fa-2255-4a52-aecf-c71753c526b4 · outbound

This paper cites Longformer: The Long-Document Transformer.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Longformer: The Long-Document Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.743592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.743592Z digest=sha256:c0dbb976299064fa7854d7d461494864a430d91e5d6a59f924beff582ec37ede

Observation dd0ba8b8-d185-4029-98eb-6479edf9cd5d · outbound

This paper cites Efficient streaming language models with attention sinks,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Efficient streaming language models with attention sinks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:17.413962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:03.876438Z digest=sha256:16aaab678d89813fd375ea2b55d6d92c64a7f4846133c998e7e00453055482c9

Observation 9c1f97f4-31be-47bf-a7e4-89aa14373d90 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.996288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.996288Z digest=sha256:f0c5c538c7fa52f5b4ab33ed4e6aa8908a62381b4cf5fb04109607e8fbcd346a

Observation 880c0f54-d757-48f2-b888-07afa42268bf · outbound

This paper cites SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:04.130346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:04.130346Z digest=sha256:d33bd3945ecfeaf9ff98cafe82bda7a805c64f0a3e04d7ebf102f445c3390efa

Observation 23c86693-7505-4dfc-ab59-2ab4c6e8a005 · outbound

This paper cites Flexprefill: A context-aware sparse attention mechanism for ef- ficient long-sequence inference,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flexprefill: A context-aware sparse attention mechanism for ef- ficient long-sequence inference,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:17.203121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:04.292807Z digest=sha256:2ad1ebb6d7150bd13195bafd7891d7b24df98fcccb7ab6befb286edae101c007

Observation c91011ad-266a-48c1-a33f-3149672e4140 · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Spargeattn: Accurate sparse attention accelerating any model inference,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:04.457185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:04.457185Z digest=sha256:6c1d926826afb8ee9421b22bb4576f109c313f0a03bad6c305a09a098bed3b76

Observation c552fb85-094a-49c9-ab38-0dc97356c943 · outbound

This paper cites A training-free sub-quadratic cost transformer model serving framework with hierarchically pruned attention,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling A training-free sub-quadratic cost transformer model serving framework with hierarchically pruned attention,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:16.819110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:04.564148Z digest=sha256:5c2fd35ae13354d7f8925a0dac6cc0fc8d2d67b0f6082ec78b6b0ed6704ac4d3

Observation 4d02b2b5-2c14-4ecb-9d69-11e44e8a1782 · outbound

This paper cites When attention sink emerges in language models: An empirical view,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling When attention sink emerges in language models: An empirical view,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:16.468229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:04.678368Z digest=sha256:cb87777eb2c942dd214b9713fe7a8392b9803ae9de0bd1eb0fd0b25b8d2d2681

Observation 777fd1ba-d746-48d0-8058-35432def5b16 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:16.209941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:04.867954Z digest=sha256:f7c0b56d34127c0390c9c7e62ac3f2306762349bd4d37e42b75b6cd3b39133ed

Observation 0b2391d7-3bc3-4e6a-9457-1abb8c98786d · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Snapkv: Llm knows what you are looking for before generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:15.903699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:05.026810Z digest=sha256:756989f3e787629c083c24af4d99fc4d9c7bb691aa9121b376ba2bf00db2767d

Observation f0d3eeb6-e310-4965-99c8-b85aa2d41096 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.152541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.152541Z digest=sha256:103e63d16125caf7ff6cde54ef50e5743f476a25b7e268bbb772a7742c14b1b0

Observation 68f5277a-7d78-4afb-813e-57c80128c8d0 · outbound

This paper cites Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.263720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.263720Z digest=sha256:e6bc7b36d8684e66d74cd6611ce421008254dad92715c36589c2e8f43f41b05b

Observation d099f893-9734-41ba-994f-43a4a313878a · outbound

This paper cites Qwen2.5 Technical Report.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.385233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.385233Z digest=sha256:f0de2644e95f8b3fe3e12f53a266f2063fe14050861db745a2c80a5a89500f73

Observation 8fbdeb21-5945-4ade-a74b-dfe82e553374 · outbound

This paper cites Characterizing Prompt Compression Methods for Long Context Inference.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Characterizing Prompt Compression Methods for Long Context Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.504381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.504381Z digest=sha256:807d0fd8ccb514d3b86bc309998e5efd3be518c8c60ad3e6d5acd4f0df154699

Observation 0e17897b-9f3b-4af4-bd78-0b15be0fe53d · outbound

This paper cites Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.601985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.601985Z digest=sha256:3fa41fe6df0ad24e8552176f75163e652b8251cada2ddcaa8a8c19b9b26309b0

Observation 7fb1ec33-d153-40f9-9093-18fd3a040d60 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.729533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.729533Z digest=sha256:7f4a63a89c82d81c6802439f75444a596cf78ae0aa11b37397152dea22376d7d

Observation 8677bee6-8834-4439-a1f5-446277b5db5b · outbound

This paper cites Compressing Context to Enhance Inference Efficiency of Large Language Models.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Compressing Context to Enhance Inference Efficiency of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.895555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.895555Z digest=sha256:bbb284b01eaededf11e78b3de2f2d5cc3b91652ab44473104c5d3c12a363671e

Observation 54ad1def-52a4-41d4-9065-0738cec69d22 · outbound

This paper cites KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.015867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.015867Z digest=sha256:b99ae8dcc05275fb75d78833160176bab48e50408889f6c01495490040bcdee3

Observation 73c9d2ff-0379-4a51-80bc-4a1411f3cd37 · outbound

This paper cites Moa: Mix- ture of sparse attention for automatic large language model compression,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Moa: Mix- ture of sparse attention for automatic large language model compression,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.153339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.153339Z digest=sha256:69a45c4db7fb8f8772b5b307c0bea59485847f173938b704444f8abfeefebbd6

Observation a9985a76-0204-4317-95d5-3abf1cd66568 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.317237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.317237Z digest=sha256:d6985216a62120263ee9c1801a8ab33fa9c1f2e9ea7161a5132a5cbe12dae2f8

Observation 783534c0-5b4b-49b8-bff0-f2b48d7627f8 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.471625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.471625Z digest=sha256:2f278ab41fda873f19b1b29e190e36c4a7a7b978cbefe914c19ea1dce96bb760

Observation cb1b6144-7239-4bab-a852-bd4d7f9f8a3e · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.633344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.633344Z digest=sha256:769da1133e557ef02338e1cbf5df3aef029c16c0c7ff71c8063823e6532e5c41

Observation 52639c41-fccf-42ba-a00f-b3047216356f · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.783759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.783759Z digest=sha256:ff22db85fb797ee1686378b6ab1334e1422e88fdb3583ccf0934a76e8ffd3936

Observation 122136bc-932d-4d23-a356-0c9231bfa3bc · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling RWKV: Reinventing RNNs for the Transformer Era

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.925763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.925763Z digest=sha256:bcf30342c690f604aa2490d441d83c32b43095818f8424121dd7f74c7bdfc498

Observation 9ffea3ab-fd1e-43dd-ac41-cd4805e84c8e · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.074760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.074760Z digest=sha256:cdcbc48e0b1679e6d7384eb5164801dee7f686cace6778aa5c1e102055232d0b

Observation 072a6031-85e0-4157-9026-1e82c3954bfd · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Mamba: Linear-time sequence modeling with selective state spaces,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:15.549430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:07.195393Z digest=sha256:f5c5994ad4a6c68097c6dfe48de7b4f2ede91afae8e8dc03f8464111b24fcba1

Observation 656dcb64-0540-4ede-b338-6e62a40d085d · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformers are ssms: generalized models and efficient algorithms through structured state space duality,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:15.233636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:07.353698Z digest=sha256:e55e89271cd1c36f5c1c17532a3395a435ac1e1d24b15f4a86b5aad70f20d8ef

Observation 8094898b-a09f-4271-ac6e-1e47a32bf967 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.482431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.482431Z digest=sha256:a2f8bf1cb8830f5348aa438673e90f3ddab9fcbb887433b6a4aca2323a3a5efd

Observation c1ede19b-d24e-4f37-9da5-4551d0f4b7eb · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.895773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:07.591420Z digest=sha256:589e30020c02ceda6dd71aedc02d02b9136c25d2df85650765082931aa89d346

Observation c8b8f53b-8397-443e-8325-0edc21cec549 · outbound

This paper cites MagicPIG: LSH sampling for efficient LLM generation,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling MagicPIG: LSH sampling for efficient LLM generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.691354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:07.803444Z digest=sha256:b5e99e9c69e6aaa6a9000280c145b1e7bec152f8cf1951a123f9e98ccf43cab0

Observation 421f902d-d766-48e5-9a7c-159250df986e · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.925816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.925816Z digest=sha256:5fa2760aa8359430360a1aec2a28ee8d429c362627c300129ccf83238463a164

Observation 05e50e07-b88e-4883-b6a4-fc715debffd4 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.514745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:08.057556Z digest=sha256:a4d6e88537f1e0c4b0d13a22ae099c8ed32714789740981996fe1593c0f0686b

Observation 815689dc-7de3-4398-8a7d-9d441cf63646 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.191015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.191015Z digest=sha256:f02ba234b5b72b640909d28ef868bddfa0f6e8d4cd08035f00bd83ef01f8ec26

Observation dbc3a49e-e94f-44ed-96cd-72c9db4abe33 · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.367652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.367652Z digest=sha256:83b4d2edbc320d076ba60c756bd5be4ca8d282542e5976b576083736c5ff38e8

Observation 7771ceee-8c80-42c4-ab8e-5e1c995482d3 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Cam: Cache merging for memory-efficient llms inference,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.395429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:08.545492Z digest=sha256:354f67ec2aae72f573d87e0788275dbc0204d220c336318837f9d3ffff3bff93

Observation cc45d404-06d6-4b1f-b8bb-d6a0af867bdd · outbound

This paper cites SubGen: Token Generation in Sublinear Time and Memory.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SubGen: Token Generation in Sublinear Time and Memory

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.668655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.668655Z digest=sha256:723a4abf6f8d1f7bc3cd1c5443dc3049f8a52a267b01f0526511769af997849b

Observation 7d3e88fc-7553-4943-bd0e-8208da72cbd5 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.199089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:08.838348Z digest=sha256:f9aac7f897d2a0548b2e9be0830d1d912ddfe262dee25235cec8acb06aa71eb9

Observation cc81e05c-8df2-4ac4-9c0f-1c127d93a221 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.999098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.999098Z digest=sha256:7603bd9c97ebc108db118de5717ebf26bf7ea18c1f3c0bbe20fdb3d2c2279621

Observation 0a37b6a5-5f58-4758-9f2b-8da82c070851 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flashattention-3: Fast and accurate attention with asynchrony and low-precision,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.945945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:09.165987Z digest=sha256:28d4ca4b90ce412811fbf8bfb551a03639eaf9aae457b6725369262b9b2a7dfd

Observation d15aa6d2-c9ba-4662-8b4f-9464d76ad745 · outbound

This paper cites Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.239572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.239572Z digest=sha256:b27df3428a5105f77b8ef8efcda16cdd5c1177ea1782923ea1d8ee02c05a858e

Observation b210d8f7-8e27-4561-b883-0e2d0f162c3c · outbound

This paper cites Sageattention2 technical report: Accurate 4 bit attention for plug-and-play inference acceleration,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Sageattention2 technical report: Accurate 4 bit attention for plug-and-play inference acceleration,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.376740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.376740Z digest=sha256:9f27d87d12ed261f7a869befa302e26e4e9fd385e1b1734a98eb8180c022d009

Observation fb3fe215-278a-4b2a-a516-3d74a2c82088 · outbound

This paper cites Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.794860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:09.493235Z digest=sha256:4b50a47b9045719dbef0047585c55cfed4b7f3d89ce4e58f620c9c3ef297b127

Observation 0a3992ac-f256-4aa1-87be-50971fe84b26 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Triton: an intermediate language and compiler for tiled neural network computations,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.573170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:09.618175Z digest=sha256:c1096de35425671a3de29242b41d1be4e8bf7da6bc584308d2eb93e485123716

Observation b2475b7e-d19d-470d-bab2-1864008f8952 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformers: State-of-the-art natural language processing,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.287003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:09.724622Z digest=sha256:c35ff932c60bd07723d17059cfd23f2ec7707ee9f7f964d7bb77854b787e7306

Observation b7f2c074-e5c4-4711-9955-6e96db117fca · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.826529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.826529Z digest=sha256:8c56070cf1e28ca34a0351fd70655e7b0437e44a32e0fa02b7cf4a01dcb14ccd

Observation b44475d5-4d5b-4b6a-aa3b-fe28cae2bbc4 · outbound

This paper cites DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.890077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.890077Z digest=sha256:b0152ea2a6193a3c0e87ff6bb3ba05cae3abe8548192753534c7473ba0a03b82

Observation 3b2b6213-5e68-4e18-8bb8-1ed454617470 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LongBench: A bilingual, multitask benchmark for long context understanding,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:12.976678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:09.973087Z digest=sha256:1242c924f130cde03371e9c4ddd528ab13de64aebf4717a3f23a7a9045848e0b

Observation 01963edd-4a23-4919-b6cd-d2123191a12a · outbound

This paper cites ∞bench: Extending long context evaluation beyond 100k tokens,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling ∞bench: Extending long context evaluation beyond 100k tokens,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:12.732486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:10.082023Z digest=sha256:699df9a815bfab755ea0df07bd3a59e29f600f142523cc50180b27cf70e8a4c4

Observation a29cb238-8b36-448d-a202-e4a809b09dea · outbound

This paper cites Needle in a haystack - pressure testing llms,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Needle in a haystack - pressure testing llms,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:12.426025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:39:10.193536Z digest=sha256:4fdbeeb97176debcbf51ce552af2c54347dc7275e08d3c7f3738a9241f574e41

Pith citing papers

No inbound Pith citation observations are available.