Pith. sign in

Paper Citation Record · LEDGER

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs

As of 11 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2502.06766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06766 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:29:13.953052Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:05:02.658992Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:31.567653Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54993fde-173f-4101-b8d6-417b5c39b4b3 · outbound

This paper cites write newline.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.759814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.759814Z digest=sha256:35b2eed7f860d7b8dff244367afb0dae7bb1def914014feaf88392ab4cb5a472

Observation 8550423d-459c-488f-9614-bd0382e27052 · outbound

This paper cites write newline.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.765443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.765443Z digest=sha256:f4d219bb0cf84ff6d21a6b68a8fa07750b927b250bfea96c5fc2cc8fcf5743ad

Observation db06c88a-287f-4fdd-9084-f482a061eccb · outbound

This paper cites @esa (Ref.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs @esa (Ref

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.769661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.769661Z digest=sha256:35a1e8633e88c282978a5f208273a9dcb365afa8512cbd95223e016c38179046

Observation c79d6a49-971f-43f5-b8a6-5b47984d16bc · outbound

This paper cites an unresolved cited work.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.773753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.773753Z digest=sha256:a99358024236bd5da8fa7b918680dd9b659e6a87509e2374abebcf147c5c3713

Observation 551d0c8d-6e99-4471-8002-921546cf9a2f · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.777883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.777883Z digest=sha256:e15bb73d1117a0afb1749a328dfaa5e7e3b1d18788c24df32a160df9a076871e

Observation f18732d6-7d8b-4d2a-a471-c5580e0c1028 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:29:14.583100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-08T14:29:13.783553Z digest=sha256:25dc61bb4765032c6a1332c7841940cda7c29679297fa263e3d6c1a81b0a3227

Observation 3e11d60e-ac7b-4254-b1ab-738c649b4437 · outbound

This paper cites Longformer: The Long-Document Transformer.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Longformer: The Long-Document Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.791771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.791771Z digest=sha256:8b0065d24cfb9921147eb57fb807da9e07d2e670e6eb3d8557002d9531ac1e23

Observation 902399f9-3fd8-4946-907c-36eaeabe1a28 · outbound

This paper cites L., Gao, J., and Choi, Y.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs L., Gao, J., and Choi, Y

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.795628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.795628Z digest=sha256:f1e289e90e053a7dbd12736513de5e1e605cf508c305179e30fbf6182af125f0

Observation 7e6550d2-bf47-4809-b6c9-89a45ac5e230 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.799090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.799090Z digest=sha256:c1e25719d60841b3120cbbd234072562f28bf6363453f858083329f7878f30c2

Observation 38ba009a-6396-495c-80c8-c86c0e8d7ed9 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Generating Long Sequences with Sparse Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.803285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.803285Z digest=sha256:9213a420b4141163dcbd71a46fc0ba6ff641f2563c7d23a9212fe86b4ce68ff0

Observation ab71d7a5-b318-4792-94e7-20df258069e6 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.807304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.807304Z digest=sha256:bcfa4592a28cf135c38ddc436515be9f56cd642209e642f60ef8c061d0ec1436

Observation a2a321e2-9e13-4b17-98f3-d55fdd816210 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.810775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.810775Z digest=sha256:661e0e977e07439aba8fdef79247aae4c795151fdc5612b846f172a40cf16b90

Observation 6fdb94af-4f89-4625-ac00-ebf00adc9c61 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.818243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.818243Z digest=sha256:87cdfbfea7ad0640f850127e0564713686cb1ea5596dec67630aa9f58ef208be

Observation dddbe6d7-8fd2-495b-a4a2-848671133341 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.822810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.822810Z digest=sha256:c7d3841e9c7415c44d9d754f62a66b2e4e3fa9d5e404dfce0543963659528466

Observation 27a19358-5f05-4083-a7b4-fef243487a6f · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.831220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.831220Z digest=sha256:76703432f1526b4b416dca986502983c702042a0797830ef29b672b1c95a04ec

Observation 256d080c-c0d0-4f3c-8638-177dff01874f · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs A framework for few-shot language model evaluation, 07 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.835668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.835668Z digest=sha256:d6b9cd239f138651e1358e63dc40c3280ccff8b69c180fb4652bd8854a66e3a6

Observation daacac68-35aa-44e6-b389-3d9ed965d1d3 · outbound

This paper cites Scaling rotational embeddings for long-context language models.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Scaling rotational embeddings for long-context language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:29:14.549210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-08T14:29:13.839482Z digest=sha256:89683ffa40bb922ea4fbcc686fc774d2c2030612a4f55e29795a52de2dde0d64

Observation 2a2af44a-adbf-479c-a628-5ca521e8099a · outbound

This paper cites The Llama 3 Herd of Models.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.843112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.843112Z digest=sha256:84f50f88e1952df4007b7ad6c32dc02e87dc764b3e67fa8539e3eac78b624b82

Observation 8278ec27-444a-4a53-b670-5bf17e096aa5 · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs The Unreasonable Ineffectiveness of the Deeper Layers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.847195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.847195Z digest=sha256:743ccf18ff14bb537f7acc90862f2acb35121e183c4ceab5e4b81a47adc09532

Observation 153b8895-84a1-400e-b7e2-27957bd983ef · outbound

This paper cites Memory-efficient Transformers via Top-k Attention.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Memory-efficient Transformers via Top-k Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.851910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.851910Z digest=sha256:4bef36990db1cf6110af50469fc496cf3394900343f21eb7abcf46052535199b

Observation d5fd0138-942e-42e8-b4ef-46d3f10192e3 · outbound

This paper cites Measuring massive multitask language understanding.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Measuring massive multitask language understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.856191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.856191Z digest=sha256:632448307396753ef1115025df4a104ee04ffd9400dc6f79ac6ef24189388428

Observation c3a16dc3-69ec-4550-aa0f-3b0b603b306b · outbound

This paper cites RULER : What 's the Real Context Size of Your Long-Context Language Models ?, April 2024.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs RULER : What 's the Real Context Size of Your Long-Context Language Models ?, April 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:29:14.530816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-08T14:29:13.860515Z digest=sha256:ce90cedba5637f641d6a56ee365e5e9689be6f421a53f5b114231ae465d8ad5e

Observation 397f8154-9083-4bad-bc71-f41efec1583b · outbound

This paper cites Needle in a haystack - pressure testing llms.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Needle in a haystack - pressure testing llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:29:14.518574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-08T14:29:13.863917Z digest=sha256:2abf91b54767d432a3b512242ad6b8d8aad7433cd9f625a295173432ae6e51bd

Observation 65abee94-256a-4d90-bd22-34aa31e4eb72 · outbound

This paper cites B., Chandra, B., and Yejin, C.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs B., Chandra, B., and Yejin, C

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.868392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.868392Z digest=sha256:c053e10fa895381fa423944cb8d373be1d0e9ddfadeecb58b3979436b825d616

Observation 84d35603-d1f4-4140-8372-d63d0a8911d4 · outbound

This paper cites H., Gonzalez, J.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs H., Gonzalez, J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.872596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.872596Z digest=sha256:bd600133d580d0286c80482dcc232ea428ff4be22398edbf64da006b8d3edf57

Observation 1635e9b8-f15f-4503-8824-331e79b3e922 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.876668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.876668Z digest=sha256:9970b009c26177c2c4daad920ef50f42379aff5722ec302aabcc3a5bf0311957

Observation a1483b56-843d-44a4-9269-22ae9cfdd3f2 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.880612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.880612Z digest=sha256:d77d36825caf8cf813e0048d6d551deb0ad1b79b844069842f8c7429fe536719

Observation 2fa32ff4-c7aa-4fb6-8223-a64abfea562b · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.884896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.884896Z digest=sha256:db5a7c4915986e2ed05da2460460ba7b4f478563d6632a638a5df627d5d1113a

Observation d5b8732a-9843-4fb1-9c48-88de3aac92b3 · outbound

This paper cites an unresolved cited work.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.888866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.888866Z digest=sha256:00fd1b6941fd0350000a2ffb1008c5e2fb245cce8241a05c699b648cf93520d6

Observation db090933-f005-432c-abaf-5fbe244166d9 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.893490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.893490Z digest=sha256:6626730c4b3d82a700b84ae87882ee5ee055d348acc86fdfa12064413f662341

Observation 44d212e6-333a-4f3b-9bc4-f5368a0d2780 · outbound

This paper cites Efficiently scaling transformer inference.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Efficiently scaling transformer inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.897639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.897639Z digest=sha256:bf6e13621acde3f35c5508670f466c1526439fb85228879343a7e55da47abaaa

Observation 47658574-3561-47a1-8750-9c044bf123e2 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.902124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.902124Z digest=sha256:104228d1aab196365ac5a7f07a395c22db3556ced841caf655b1c5219989368e

Observation e383ab97-b82b-4b0c-b505-222d5de198c5 · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Flexgen: high-throughput generative inference of large language models with a single gpu

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:29:14.478749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-08T14:29:13.906745Z digest=sha256:cbbaed40700d42e31d9a6cc8d054b1fe053d4fd5eb637fb9e50d2e368870ec8c

Observation 4930a632-4f1d-42cd-b2d2-53356ebee0ba · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Loki: Low-rank Keys for Efficient Sparse Attention

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.910553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.910553Z digest=sha256:3b7e5832f6004a32a19b9c59d6eb799ded9c62450cb86ff6ed5ac1134e77f5f6

Observation 1fae160c-2723-4657-929d-efd23f51e6a7 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.914777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.914777Z digest=sha256:f38d64ab844305cd20b65e401291b143bdb781c4754ee8a65511ca4b9d4f2440

Observation 691e681e-023f-4a2c-b3e1-c08b26c8532d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.919045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.919045Z digest=sha256:d5cb40a4956b7f78adce96981ba6133f511617fd3d8c44f1f70adeb421cb4912

Observation d000197d-d80a-4e49-8620-e2a18b2a19cc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.923421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.923421Z digest=sha256:870570c0ee55a395f34a4ebbec65acd58ae177f3dd496cf1a35224ec9a9e34e4

Observation b8f6d6e6-410f-4b73-bd07-0baf78eb830b · outbound

This paper cites Efficient streaming language models with attention sinks.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Efficient streaming language models with attention sinks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.927821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.927821Z digest=sha256:7776ef48b5a39007917f355768a42acc5d7252915f35a2a86c9e51db5f29f078

Observation acdc13a7-f186-4d14-8138-685e4382e4ec · outbound

This paper cites Efficient streaming language models with attention sinks.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Efficient streaming language models with attention sinks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:29:14.461596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-08T14:29:13.932114Z digest=sha256:258faed50a704329e670dbe7423789bbce689f10cf345fcc3002256911152714

Observation 8b5e6642-90a0-4b57-baf0-35762383ff09 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.935794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.935794Z digest=sha256:3ca6538a5a58e6f577d6f02814cfd1d2fdf3ddbb3f26a7441e7aa3b516c16f2b

Observation c66cce35-626a-4433-b689-b2d630ad696b · outbound

This paper cites an unresolved cited work.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.940205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.940205Z digest=sha256:5a765fed32d09c103b0e8901704a1a50fa419fea833f77345615ae253fa07705

Observation a0e7f2bf-fb87-479a-8ca3-3ca35c60e3a2 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.944332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.944332Z digest=sha256:ab4b3ecdae16c353562753bae89320853c65290319de8e1a4111c753565c3e10

Observation 78d04669-c49c-493f-94d2-50c292f429da · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.948667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.948667Z digest=sha256:bfbeb78e6d03fb0ffcde9300299d35b62c0b6c978dad6b6a0a55d27a0d208e73

Observation 8e949884-6027-4bbc-bfea-653ac3a2ae1a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.953052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.953052Z digest=sha256:b83d88b856587c2882b6a3976bc8357a11a287fe427fa0d6acdcc0e93c1c6bfb

Pith citing papers

Observation 0048375a-034c-445f-aa2a-f55b6ab1c697 · inbound

Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs cites this paper.

Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:59:36.902216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:59:36.902216Z digest=sha256:977769609e94887c6b511250e6f0df4b5cd9a18c624e32b3004c183c1d131c06

Observation 57282a15-f411-456a-84a5-032af00dde3f · inbound

Attention's forward pass and Frank-Wolfe cites this paper.

Attention's forward pass and Frank-Wolfe Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:05:02.658992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:05:02.658992Z digest=sha256:277e6095b6ce83e5e213699ca3bb558f61bb41191c1fd61f6df78c8ed24590b3

Observation 2e674d1c-95a9-407e-bb2b-6644d511d0b7 · inbound

End-to-End Context Compression at Scale cites this paper.

End-to-End Context Compression at Scale Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.569092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T16:36:54.699174Z digest=sha256:e82975dd38e57b1c352b744fc376529eb81bdbd9e717bd41e9aae7ff0a01823c

Observation dff7b74b-a10c-4dba-bbb3-fdc41ad26aed · inbound

LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention cites this paper.

LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T07:07:39.861302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:07:39.861302Z digest=sha256:4008866ed985599d10f541e65a527011d8884e76fce22628ba76dd00827a02c2