Pith. sign in

Paper Citation Record · LEDGER

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management

As of 22 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.07009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07009 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:45:23.455170Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08450613-eb6a-43d6-aae1-f173b5525e8e · outbound

This paper cites LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.291458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.291458Z digest=sha256:11f80ec49d753895ffc65667659b1d9aae9b98f5f64488dd02d9e50fb83bdb6d

Observation 1b827acd-2743-44e3-9406-16faa4e31499 · outbound

This paper cites IndexCache: Accelerating sparse attention via cross-layer index reuse, 2026.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management IndexCache: Accelerating sparse attention via cross-layer index reuse, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.295722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.295722Z digest=sha256:4feee27a5b0b85307b12d6613fe41c82efa844a2c9cff447dae2c2d4974d04eb

Observation 18eecb4a-35bc-4e38-9ce3-169bcea1ace4 · outbound

This paper cites an unresolved cited work.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.299171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.299171Z digest=sha256:2e7e3338cf6fa93fe221ded333a7586bbd6e110867df9d585fd7d1f2531ccbfc

Observation 63cddaf1-c477-4db0-9da6-49064ca8accf · outbound

This paper cites Peters, and Arman Cohan.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Peters, and Arman Cohan

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.302928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.302928Z digest=sha256:0dbfd0e3dfe837b4539a5c25eb5702a22df9729337adc36b5cf1a2c96e91c65c

Observation f4fe85e3-d7ab-442d-8289-be80dff3b30e · outbound

This paper cites ArkVale: Efficient generative LLM inference with recallable key-value eviction.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ArkVale: Efficient generative LLM inference with recallable key-value eviction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.178841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.310628Z digest=sha256:7255c5ab2eff7a2b860e1043e3b1a3151e8b027866ed47b3b9a9f8b1c0186bf9

Observation 0dd0d590-d780-4eb8-922c-7edb139f8458 · outbound

This paper cites ESS: An offload- centric latent-cache management architecture for DeepSeek-V3.2-Exp, 2025.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ESS: An offload- centric latent-cache management architecture for DeepSeek-V3.2-Exp, 2025

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-10T16:45:23.810177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.314548Z digest=sha256:8498702c733a8fe505965f5bd9eb1616da90fb6f2d76d44869a8189763427c82

Observation 206a681e-7d02-48f3-a697-3ce2dc7cc6f7 · outbound

This paper cites MagicPIG: LSH sampling for efficient LLM generation.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management MagicPIG: LSH sampling for efficient LLM generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.167750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.318523Z digest=sha256:850cf09ee0a9083ca5bf944cd4d9b6fa44a7dc19af138453135d00c8b5ef6669

Observation bf2644b7-59bd-4cd7-9575-b6aee5ef0486 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.156172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.321833Z digest=sha256:69160cecb400b87c09ba1b5263bcb4a04f9d6e400d8a1ef4d4f49d3d78014415

Observation 443f7eff-8486-4d41-9538-da640552c794 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.145252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.325223Z digest=sha256:3b6f28b4b5607596076b99721739543ef96b01dc393ef20b2b4d6ae509cc3e97

Observation 950f44f3-a46f-4438-a616-2d5359620b94 · outbound

This paper cites DeepSeek-V3.2: Efficient reasoning & agentic AI.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V3.2: Efficient reasoning & agentic AI

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.132858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.328699Z digest=sha256:81b040c297a29456a9c00189168f96b9bf0c74ed46b4fa60540d0b701bd3f2da

Observation 51644309-8908-4f0b-8327-d9b009988433 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.332283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.332283Z digest=sha256:c4d314784184def4a01856891b3265a3ca2db0f5c65e4e9120ef612ffce2c15e

Observation 87e82dd7-d097-4c32-a51c-b445c0fcbad2 · outbound

This paper cites DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.335996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.335996Z digest=sha256:a92fb100cdda28ea2c8a13f7da5526410c2443e0f69109d01fd8606cbf48d474

Observation 6639af0c-5fa4-4eca-b8e8-d67202168ffe · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with CachedAttention.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Cost-efficient large language model serving for multi-turn conversations with CachedAttention

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.120914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.339727Z digest=sha256:d1cffac943e4a4fbd77a07f0f7dfe46e4a477382f8f0b7024d4e696325f7442f

Observation db37646f-9abf-42aa-b2a3-918176736f6b · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management GLM-5: from Vibe Coding to Agentic Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.343078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.343078Z digest=sha256:a85842c867f07cfb5efaec2a03775d48862a7357886035588d3ddd3768de831f

Observation b101c3db-c662-4e7b-aa80-9a0f26e7f45c · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.346750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.346750Z digest=sha256:a897d86e45a1fdfef6fc876f4ca4654dfee364ee8dc77f578b15223a76717fa1

Observation 617066a6-f89a-486d-afbc-84989ba4f922 · outbound

This paper cites NEO: Saving GPU memory crisis with CPU offloading for online LLM inference.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management NEO: Saving GPU memory crisis with CPU offloading for online LLM inference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.109280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.350702Z digest=sha256:2c24cc8b830eebcaacf96f00f4a594d10c53ea36cd12d0b87a82b7ea20aa4eb1

Observation 95365699-59c1-4ad6-8de1-4cceda3b1a08 · outbound

This paper cites Reformer: The efficient transformer.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Reformer: The efficient transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.354354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.354354Z digest=sha256:d3f1c38db1161924af46a2a0f4f118e6ae58c493c16a66be7dd6def2605f8deb

Observation 7b6669bd-e03c-467f-99d5-3191d56b2bbe · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.358099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.358099Z digest=sha256:a527b4cbf24ecf104401bf3491f74b6c1d713047b4bc2501a1ddf4662dc47785

Observation c3219acc-acf7-418a-98fd-1dfa903ad32d · outbound

This paper cites InfiniGen: Efficient generative inference of large language models with dynamic KV cache management.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management InfiniGen: Efficient generative inference of large language models with dynamic KV cache management

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.091321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.361753Z digest=sha256:7b046e5f8ca9bcf4900f7face8afd27238eb7bae273527627470cdf4f7f224cd

Observation 037b387f-a27f-4a2e-85ad-31ae52571801 · outbound

This paper cites SnapKV: LLM knows what you are looking for before generation.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management SnapKV: LLM knows what you are looking for before generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.079545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.365570Z digest=sha256:94ca0beea8c4735cdbc7accecee3956ec7f3da8e04733d00299421c2542be856

Observation a899c11b-847a-401d-ab04-7ac5315e7d33 · outbound

This paper cites ECHO: Efficient KV cache offloading with lossless prefetching for serving native sparse attention LLMs.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ECHO: Efficient KV cache offloading with lossless prefetching for serving native sparse attention LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.068201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.368982Z digest=sha256:b3dd5be065aaaab77592d18c8a89b8eba543276d41a6cc22d505e8817fdef056

Observation 18955ac0-7ab2-4126-9599-a8b5f22a4dc1 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.057458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.372519Z digest=sha256:4b780c8f992487b6b21c989caaf301ae334b7e3b72ca227e261cac9763e82599

Observation e5d2f803-6b22-43b3-9283-ec291a443307 · outbound

This paper cites NVIDIA GH200 Grace Hopper superchip.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management NVIDIA GH200 Grace Hopper superchip

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.046620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.375879Z digest=sha256:eaf0a02b3ad2fceb5eb909024415a98324c0b5d0979b422d8ff3fc0be84fc115

Observation 1b2b2c5b-985b-4edd-acd9-cb1afae4eca1 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Splitwise: Efficient generative LLM inference using phase splitting

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.034020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.379612Z digest=sha256:4cd3d74103a9d6591906f87a047b697082531388ae2882585e1295f7fad1c568

Observation 2c8afc0a-6b0f-4cb4-bfab-05e645e472c4 · outbound

This paper cites Mooncake: Trading more storage for less computation—a KVCache-centric architecture for serving LLM chatbot.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Mooncake: Trading more storage for less computation—a KVCache-centric architecture for serving LLM chatbot

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.023147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.387043Z digest=sha256:1f306dd039c1878b249e7765cdb6ccc6cfb2eea9552327fe735a43909a26c4b0

Observation 47246dca-dccc-402b-9779-3fcad698a246 · outbound

This paper cites Qwen3-30B-A3B-Thinking-2507.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Qwen3-30B-A3B-Thinking-2507

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.011712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.390523Z digest=sha256:fac16365ddb564b6f379e15cfe4929eeb326f207fe1c0bd300da95ae3030dd4d

Observation 39509634-85f7-41c8-a6c7-7517926305b4 · outbound

This paper cites Bench serving guide.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Bench serving guide

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.999816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.393999Z digest=sha256:1f86c4057211d2be89b707a40c3fd6ceae0dbe850ed536877d7c531642942b24

Observation 4b0eecd4-41e1-4e3b-8ef8-a7534b2d602c · outbound

This paper cites HiSparse: Hierarchical sparse attention.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management HiSparse: Hierarchical sparse attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.987113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.397212Z digest=sha256:9a4991a4ed82b7a71a1b2000e089a13bdffa166f447b533197a3a656ab59b024

Observation d841648c-33c0-4963-9330-7a92637bf412 · outbound

This paper cites FlexGen: High-throughput generative inference of large language models with a single GPU.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FlexGen: High-throughput generative inference of large language models with a single GPU

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.964365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.403797Z digest=sha256:8db50a3e9cfc9012c58ed81367d927de08c6dac43b6f2717a777ca1cd2591668

Observation cbefa9c6-2a90-43ae-bb7a-7eedeebdf0c6 · outbound

This paper cites ShadowKV: KV cache in shadows for high-throughput long-context LLM inference.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ShadowKV: KV cache in shadows for high-throughput long-context LLM inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.952778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.406916Z digest=sha256:821de20a5064d728817739cd97e51132297982053e4bcfc5d247d2dcda981fde

Observation 2ff478fc-ffba-4d28-be6f-4d8d73867a95 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context LLM inference.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Quest: Query-aware sparsity for efficient long-context LLM inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.940899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.410330Z digest=sha256:84fda018dd76b21e032ad2916b9cc59321a0fb262c990b9f8ad6c45d800fd7bd

Observation 610d0934-c4f7-4171-906e-006158d27888 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Linformer: Self-Attention with Linear Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.414556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.414556Z digest=sha256:35837f318d17ae59640f2d0b3d2759189f7caae4af5a40748b3b93a0c84d51a9

Observation b2f6430f-d3a9-4593-8931-bccccdfc123d · outbound

This paper cites Efficient streaming language models with attention sinks.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Efficient streaming language models with attention sinks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.419073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.419073Z digest=sha256:9493437140a05163d5cf9024c749130aca4585facaae22e19c9df17ec8527c34

Observation 73ef9abf-fc94-4a59-9ae7-35d9ac0249d9 · outbound

This paper cites SGLang HiCache: Fast hierarchical KV caching with your fa- vorite storage backends.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management SGLang HiCache: Fast hierarchical KV caching with your fa- vorite storage backends

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.924134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.423045Z digest=sha256:a4c299dd603d8a0e1477a70e793e2bde8a4f60a83ea0e9e06d70e847b9a81a0c

Observation 229d2781-e7f0-422d-9803-08c083ce68a3 · outbound

This paper cites HiSparse: Turbocharging sparse attention with hierarchical memory.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management HiSparse: Turbocharging sparse attention with hierarchical memory

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.910968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.426667Z digest=sha256:a557f9c64f7c7ee52bbf57b8ed471e33699858692620f26bf519c6e6ff29b9b0

Observation cf222180-b25d-4802-9352-9aabc53fcfa9 · outbound

This paper cites Strata: Hierarchical Context Caching for Long Context Language Model Serving.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Strata: Hierarchical Context Caching for Long Context Language Model Serving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.429982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.429982Z digest=sha256:4a1439e24e9e8a71e157c7d2d09f51bb3b3da0faccccf52cbd915bac7f282658

Observation d4644d1c-4603-434f-9620-4e0d0f09fdcb · outbound

This paper cites Qwen3 Technical Report.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Qwen3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.433719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.433719Z digest=sha256:08f9296313abf8df6c0be1e2476767392dc4270b05a94c3e081036cf01631474

Observation 78ddc875-30a7-4cce-966b-3257618e3929 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.437624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.437624Z digest=sha256:3c8ff6aa529ad56167cd84e69fea1014412994af3393321aedb2483965c1dbc0

Observation 2c6683b0-4b82-4e1d-8838-331b0e82d336 · outbound

This paper cites GLM-5.2: Built for long-horizon tasks.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management GLM-5.2: Built for long-horizon tasks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.899561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.441276Z digest=sha256:a38310f09ec5030b4870cd2de5c1cde3ce0d76f8c4f6355679d43cd6560e1a21

Observation 126ae745-301c-44e3-8cb1-1c20ab255f32 · outbound

This paper cites PQCache: Product quantization-based KVCache for long context LLM inference.Proceedings of the ACM on Management of Data, 3(3):201:1–201:30, 2025.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management PQCache: Product quantization-based KVCache for long context LLM inference.Proceedings of the ACM on Management of Data, 3(3):201:1–201:30, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.444854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.444854Z digest=sha256:4e8ff41790d1a5a015500c615177d716dd05dd28c6ac56dbf9f5ab984c220243

Observation dbdb14c9-42e4-49cf-beff-ad7264c493c0 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.448711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.448711Z digest=sha256:c9d46e8df0c45644f795daae3f6333f7a31568cd55e057cda94dbce9d38e2194

Observation 1b4e21e5-f645-416c-a7cb-0cdc26cc8fa7 · outbound

This paper cites Gonzalez, Clark W.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Gonzalez, Clark W

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.882352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.451960Z digest=sha256:988da29a4fe96c5594566636ca3e1b647acb538211a42f81e9190ae85295b767

Observation d6f3142d-bdb0-409f-9231-eb31c460ab82 · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.870987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.455170Z digest=sha256:c2993e5470d16c56cb343be4fe824270dbf9f796fc8bef67963cc31228dc4060

Observation 7ea78599-617f-4bfd-82cd-0322b5c4281c · outbound

This paper cites Longformer: The Long-Document Transformer.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.306275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.306275Z digest=sha256:c7ca7505f7c953a81751630d5ad23ed3e04b956cdb83a3afe92b4e5b6741ccc2

Observation ccbcf82a-b759-4eb3-8e41-62a5b8a1444d · outbound

This paper cites an unresolved cited work.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.383477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.383477Z digest=sha256:ad6a6d236d70417f5a48b7d75dbba7833fe2fa7bfcc5e479ebddb2fba780b379

Observation d260293f-3b99-47e9-909d-a75a55b1efae · outbound

This paper cites Accessed 2026-05-04.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Accessed 2026-05-04

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.975469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T16:45:23.400534Z digest=sha256:6a246bb0c59009a479a1e0c9d540661890ca339c63e3c8c696b810d44f27d234

Pith citing papers

No inbound Pith citation observations are available.