Pith. sign in

Paper Citation Record · LEDGER

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.19823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19823 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:06:38.022879Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee6a794a-a478-4d69-b76b-a1f91e63179e · outbound

This paper cites 2 In the approximate attention score computation, the computation is performed group-wise.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs 2 In the approximate attention score computation, the computation is performed group-wise

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:41.183917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.876858Z digest=sha256:d74dcb165e7c6d591206270a1dfd7dc2a53074d38c5914271a6a3825f4657c82

Observation 0414ced1-bf2c-4593-b73d-89948ed9b6be · outbound

This paper cites needles" on the task performance. We take “The best thing to do in Paris is buy a fresh croissant and lounge by the Seine at twilight.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs needles" on the task performance. We take “The best thing to do in Paris is buy a fresh croissant and lounge by the Seine at twilight

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T14:06:40.923973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.880928Z digest=sha256:6371fe0746cc3d0e1c4d008bf1b22ddb904bacb1c3d53028721a4ff5b2ac8e17

Observation 96117a3d-8e7e-4b5d-ae64-ca5412b25625 · outbound

This paper cites GPT-4 Technical Report.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.884942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.884942Z digest=sha256:737a530492a426fc49a957a2bbc8821dd8b2894284458d57f2e5992ae73d5d4e

Observation fdebe3d6-067d-43c2-837a-d187f24ae57a · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Keyformer: Kv cache reduction through key tokens selection for efficient generative inference

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.739676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.887582Z digest=sha256:2cd8b7c85c2184794ff77e12eb4f83e676c4bfb45ef814ce1b04fd44b2822c8f

Observation 606dceea-355c-4ae0-bf69-50c93b47d171 · outbound

This paper cites Qwen Technical Report.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.890090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.890090Z digest=sha256:fe9820be3615b3d4c480e134d75901bab48ba618291080c5551eadfe732ef6c3

Observation ecfac449-9f5c-462e-a6f0-a0adf0e10d7e · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Under- standing.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs LongBench: A Bilingual, Multitask Benchmark for Long Context Under- standing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.465957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.892704Z digest=sha256:39b30b18b0eaaf310c3424299280e11772adf1c679f6dc7d93cd66a52819fd1d

Observation 8130c8ad-6e46-4174-b1ca-29ef63735ca7 · outbound

This paper cites Longformer: The Long-Document Transformer.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.895942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.895942Z digest=sha256:0ade6cc4f1a1e17f53127dff7ca320f319f6a42a133e0acd11a2e5d6acce9c54

Observation 389002f8-0351-49b5-8c60-8f50422b2ee5 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.235668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.898719Z digest=sha256:96fc35ce6473d0f8bd1bd2a4553fdd4c1ef8bf2231f42a97a8728ebc49cb72d5

Observation 1a40cc01-238f-4c1e-9b65-cfe96b455c10 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.901469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.901469Z digest=sha256:52388db9831bac1651abb7a9e70f87964e56cd1a3ee02a3033016d1bdb1745bf

Observation adfc8e26-5816-431e-8d23-388877f32cbe · outbound

This paper cites The Llama 3 Herd of Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.904247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.904247Z digest=sha256:151b5cefc714e15e7e559bc1802394e4ad731ac1e3e8bd6b6e7fc39858009e36

Observation 950360d6-18a2-4927-9922-c48f3179ece1 · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.907873Z digest=sha256:e0e10dcc9452fda08e201ca94c4f42d52dd97b21917201a38fd8c8bacec6d94a

Observation cca51106-df38-4722-a13a-e74117eb2f5a · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.051532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.910383Z digest=sha256:233e39a84e0b01804709040c66464b6ce3be35c6867c72d849482ee716e4cd5a

Observation 4e7b22e9-2adb-4006-a7d2-5f6170baf825 · outbound

This paper cites Mixtral of Experts.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Mixtral of Experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.917505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.917505Z digest=sha256:fa798797ce44c4ca666f6656411b3d1b1117b0c0bb7883727fb4a425e1a6930f

Observation 5fc44e7c-d87c-46d0-83a0-a26cfd1a7dce · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.921530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.921530Z digest=sha256:36eb6641b96189e1add238b647f599b6dc1bd811434af165a7a2ad23711073a5

Observation 615aef19-ece2-4278-b6c8-72c8a1b22de4 · outbound

This paper cites Kamradt.Llmtest_needleinahaystack: Doing simple retrieval from llm models at vari- ous context lengths to measure accuracy.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Kamradt.Llmtest_needleinahaystack: Doing simple retrieval from llm models at vari- ous context lengths to measure accuracy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.903944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.924853Z digest=sha256:17d41f017c0fa381c3c4257e4025a9290d2de1ad6bf07040ee30dd8a50b9c264

Observation c4c7c256-85e4-49b5-bfb6-54a9f96df94b · outbound

This paper cites Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.929141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.929141Z digest=sha256:6071bf76f20111c91ace90a7425118858d8ac4e7b950a6bdda87d6a62886f264

Observation 540ef15f-0219-466f-b0bd-cb6a60457dd6 · outbound

This paper cites FocusLLM: Precise Understanding of Long Context by Dynamic Condensing.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FocusLLM: Precise Understanding of Long Context by Dynamic Condensing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:06:38.215398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.933299Z digest=sha256:4e507c6bb18b585a209154f0032b078b91bb8a7eafba3f6336fa70dd95b25ca5

Observation b1537645-a823-4aaf-aae9-821058aae735 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.937463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.937463Z digest=sha256:1117f8fcfc0955357a4e925dfa8883cb7cf04fdea846ec504d83a45edcbc3249

Observation 2f1c78db-f234-4b0d-b048-757ca3fc22ec · outbound

This paper cites KIVI: a tuning- free asymmetric 2bit quantization for KV cache.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs KIVI: a tuning- free asymmetric 2bit quantization for KV cache

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.751585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.941514Z digest=sha256:1a9e6ea48a15509ba26a10c7f3fb82294834ea89257d6a496f99a4af6bf9632d

Observation 288d86ab-58a8-495f-a228-5b9ee5ec114d · outbound

This paper cites LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:06:38.123954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.947046Z digest=sha256:bcca4d498a4d86567e676de4c29e777c0582745d16fad5421ab04ed3b15e24d1

Observation db751f23-2830-4e43-b0c2-665b4dc94514 · outbound

This paper cites Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.951435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.951435Z digest=sha256:ce031791eab3234ed92151e14a855ed145901b6e1960b0af9e5cd78d61b731bc

Observation e5083c6d-33ab-49cf-825d-b81d20220c1c · outbound

This paper cites Transformers are Multi-State RNNs.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Transformers are Multi-State RNNs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.609049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.954862Z digest=sha256:ef4b589c1d8419a35152ebf237a1bf5ed54a2badf09dea692fd7a70b0e25a38c

Observation cf1c9358-62f9-4a64-bdc7-b1ac7d5880f6 · outbound

This paper cites Scikit-learn: Machine learning in Python.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Scikit-learn: Machine learning in Python

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.500139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.958718Z digest=sha256:e47211f72ece5375b8541c8c9800e6eeac97c8a39936342ad7f2574056a5c4d7

Observation 93f5acea-eed0-44a1-a7f2-d3b38187eb1b · outbound

This paper cites Web-scale k-means clustering.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Web-scale k-means clustering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.418351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.961351Z digest=sha256:ba205338b5115f48e9f3396a9a3cbabd2ae67d6a7901bc0bf4ed5b28189485ee

Observation acf8426f-15ab-4ec1-bc9f-acf90e0f47f7 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Adafactor: Adaptive learning rates with sublinear memory cost

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.338961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.963935Z digest=sha256:587dd1d3bd85d8a41381d2155d31db5312eae164ec25bdcc4e251b4cffaa35be

Observation d8d2ff7a-a944-4ac6-88d1-a65a44d82955 · outbound

This paper cites QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.207679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.966289Z digest=sha256:52631d33e31715f022929c358d05f41e409c45422e00e5a7abc52f8a142da909

Observation 9b79bf44-ecb7-4a4f-b923-8ee1d1742da8 · outbound

This paper cites AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer- Wise Asymmetric Quantization Configurations.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer- Wise Asymmetric Quantization Configurations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.135861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.970390Z digest=sha256:10178e5cbfcfce3f3ece8dd0a81715b457abffdd2e37714019bd28a36a753660

Observation 4a804256-821c-4af7-941a-011ea97cf392 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.974660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.974660Z digest=sha256:2c6dea7df4cd05b2b02f3a852775b16155b3d210cfe43bbbe6a7611a2732ded5

Observation 4b2a5842-f14a-4cf0-827c-9eb9dfbe3fb6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.978694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.978694Z digest=sha256:3a201e76659ff69881506f935a2ddefadbbac7970f233ce00ee1ffc18951f06c

Observation 11c6413d-c09a-4c77-9d51-63c2b6c1ab0a · outbound

This paper cites Attention is all you need.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Attention is all you need

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.904377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.981698Z digest=sha256:d3432339da9694399a04f662a22e1e0800d263e63830441f8c23ca3e450d5e5f

Observation 98f37da9-ea15-4352-a0a8-bdd8e46ac2f7 · outbound

This paper cites Fast transformers with clustered attention.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Fast transformers with clustered attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.751720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.984746Z digest=sha256:925f1a90e942e446fee147a6d8b4be793c3946b34df8af396e20e9499b2d997b

Observation 0f0aea8d-7a86-400a-859c-ac0a51eaf6ae · outbound

This paper cites SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.630982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.988527Z digest=sha256:6321a09951ce1239aaa91e4bf4685ccf69b93237a7dc06ed49da56aaaaf34bb0

Observation 1c1e24bb-8c80-462b-a0a0-6e64696e4b37 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.990797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.990797Z digest=sha256:0103e6db01884507aeee83e1d7a2f9aaab2a219787262058f2a16a9966adc312

Observation a73e0ec4-c5a0-47a7-8d9a-0825d23228ff · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Efficient Streaming Language Models with Attention Sinks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.560300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:37.993929Z digest=sha256:dc40f9adf92fda39c1b6743232ce55b9b533deec320b32b15bb26bef005b2aad

Observation 58d002b9-2bf8-478d-9a61-9193fcab74e8 · outbound

This paper cites vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.998948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.998948Z digest=sha256:85b34cc9166c97ba5f184924aea35b82d18ad2df0e002d50bfbb00173ef77227

Observation 5e20b226-1612-4e36-adc7-a2a165d53f6b · outbound

This paper cites ThinK: Thinner Key Cache by Query-Driven Pruning.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.003459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.003459Z digest=sha256:3eb37ddd7471ab3c4fe098d40f4ad3138e481cabc409f73db9ebc1939395862e

Observation c31314c2-8d44-479e-893e-6badf21f3034 · outbound

This paper cites A Survey on Multi-Turn Interaction Capabilities of Large Language Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs A Survey on Multi-Turn Interaction Capabilities of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.007767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.007767Z digest=sha256:1187e44978771b5858fd5d9a11d4bbc10432158bcd69b86c7db7784b2d0a3a90

Observation 4fe2a94b-3a14-4169-a780-0272201307f5 · outbound

This paper cites Long Context Compression with Activation Beacon.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Long Context Compression with Activation Beacon

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.011516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.011516Z digest=sha256:265f9bd62ed61d1bc8a789ff49b225ab18fbee5b6d0cc8cd6c91b07428f634b3

Observation f2ae781a-3c43-4df4-9adf-74184f70c325 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs OPT: Open Pre-trained Transformer Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.014118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.014118Z digest=sha256:b7fba73b1e41c8ebdff17b93c9bfd3828ba59758196ece248dd91bfd2a61f794

Observation 31fd5540-f7d2-41fc-9bd2-b3a1764fdc57 · outbound

This paper cites KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.481242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:38.017818Z digest=sha256:160fc8e78a3b4e1389587de52d28a972de1df02094ed935682e0fb9bc39362e0

Observation 7cb6c8ec-a7bd-4690-8ba2-937a9b53679a · outbound

This paper cites Chain of agents: Large language models collaborating on long-context tasks.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Chain of agents: Large language models collaborating on long-context tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.392823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:38.020684Z digest=sha256:497c3ed0518de4587095b7602d217132e2c7570d274a6b450b8b22d974bb0486

Observation 03539e17-272c-4263-9570-932b78338fe3 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.330399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:06:38.022879Z digest=sha256:0713507ad03f9d3c76551953ab8905cf6c3c90c2a9e18c507aa87891428924a1

Pith citing papers

No inbound Pith citation observations are available.