Pith. sign in

Paper Citation Record · LEDGER

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.19823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19823 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:06:38.022879Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee6a794a-a478-4d69-b76b-a1f91e63179e · outbound

This paper cites 2 In the approximate attention score computation, the computation is performed group-wise.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs 2 In the approximate attention score computation, the computation is performed group-wise

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:41.183917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.876858Z digest=sha256:8f95e99e95edeb1df8070935e4e69849251ad643c49b9ef70c72b764f34aac60

Observation 0414ced1-bf2c-4593-b73d-89948ed9b6be · outbound

This paper cites needles" on the task performance. We take “The best thing to do in Paris is buy a fresh croissant and lounge by the Seine at twilight.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs needles" on the task performance. We take “The best thing to do in Paris is buy a fresh croissant and lounge by the Seine at twilight

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T14:06:40.923973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.880928Z digest=sha256:cf8d4cbfd2290276652716165600082a489a9bf3e620e4a8d9b8ecab0d5d703a

Observation 96117a3d-8e7e-4b5d-ae64-ca5412b25625 · outbound

This paper cites GPT-4 Technical Report.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.884942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.884942Z digest=sha256:34801177b7a9165dc08a6c471600e6fc993d042c4c690d897f526879283c0d72

Observation fdebe3d6-067d-43c2-837a-d187f24ae57a · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Keyformer: Kv cache reduction through key tokens selection for efficient generative inference

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.739676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.887582Z digest=sha256:71326fca1bc33a5fd2f0ef48aec56f6a28aafbabf7bcafe961166d51d6b5898d

Observation 606dceea-355c-4ae0-bf69-50c93b47d171 · outbound

This paper cites Qwen Technical Report.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.890090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.890090Z digest=sha256:6b5f0cdcfc8d30796e7a045022c9a16b289a6272cfe1412ee7be08d32d38abbc

Observation ecfac449-9f5c-462e-a6f0-a0adf0e10d7e · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Under- standing.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs LongBench: A Bilingual, Multitask Benchmark for Long Context Under- standing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.465957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.892704Z digest=sha256:0cb168f4a460535173a484eeb56c638fe6ed0bb4d8134f49f24de84fc8f79379

Observation 8130c8ad-6e46-4174-b1ca-29ef63735ca7 · outbound

This paper cites Longformer: The Long-Document Transformer.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.895942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.895942Z digest=sha256:918876473b071675276e31f4118e8b8bd0422ba81dd1462ae484a5c18970a735

Observation 389002f8-0351-49b5-8c60-8f50422b2ee5 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.235668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.898719Z digest=sha256:a0b82c10477400b13817de3dad1eaf1b927d2f6f03595ecb3dbdadeeecd7b39b

Observation 1a40cc01-238f-4c1e-9b65-cfe96b455c10 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.901469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.901469Z digest=sha256:35e05776fff03adf123c3c1e12b89d606ba18551cc429d9a296ddd4b7beb2755

Observation adfc8e26-5816-431e-8d23-388877f32cbe · outbound

This paper cites The Llama 3 Herd of Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.904247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.904247Z digest=sha256:f921767990a19d9d90a6a330ee8a9efecc4710588c7c9f754c25cd174c2c8f9b

Observation 950360d6-18a2-4927-9922-c48f3179ece1 · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.907873Z digest=sha256:f7883bab441e51c412f463ee6f2d56da8869807d0ff16da5b51df1e8ead38956

Observation cca51106-df38-4722-a13a-e74117eb2f5a · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:40.051532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.910383Z digest=sha256:9d8edef4e5c6c5b45cc124b42ea79e123ca810c46703521bb00d945cb9fd2fdc

Observation 4e7b22e9-2adb-4006-a7d2-5f6170baf825 · outbound

This paper cites Mixtral of Experts.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Mixtral of Experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.917505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.917505Z digest=sha256:32fc21146f1ed44d941733b558b54e6ba14c878c00608e4761cd53290948d26a

Observation 5fc44e7c-d87c-46d0-83a0-a26cfd1a7dce · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.921530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.921530Z digest=sha256:6e0b4d97e2a275fd6e173b928f7cbcb7bd40360986ee9ca14b5066cf928d9582

Observation 615aef19-ece2-4278-b6c8-72c8a1b22de4 · outbound

This paper cites Kamradt.Llmtest_needleinahaystack: Doing simple retrieval from llm models at vari- ous context lengths to measure accuracy.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Kamradt.Llmtest_needleinahaystack: Doing simple retrieval from llm models at vari- ous context lengths to measure accuracy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.903944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.924853Z digest=sha256:971a240b3068d72f4e18333dc1a18ae9c0b79a15dc0e0664e2e17739822f292e

Observation c4c7c256-85e4-49b5-bfb6-54a9f96df94b · outbound

This paper cites Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.929141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.929141Z digest=sha256:ba67caa2453356fa401b38eb20308c2e23e2660c4f2f05caca2a694062b6ed42

Observation 540ef15f-0219-466f-b0bd-cb6a60457dd6 · outbound

This paper cites FocusLLM: Precise Understanding of Long Context by Dynamic Condensing.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FocusLLM: Precise Understanding of Long Context by Dynamic Condensing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:06:38.215398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.933299Z digest=sha256:bbaf616273a9d08dea45e3ed2a1378950d232bed765f654daf9d1872dce45768

Observation b1537645-a823-4aaf-aae9-821058aae735 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.937463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.937463Z digest=sha256:79310e793b9eff970f6b0dd449caf37587f90c54bf404674779ab5f8aa879d12

Observation 2f1c78db-f234-4b0d-b048-757ca3fc22ec · outbound

This paper cites KIVI: a tuning- free asymmetric 2bit quantization for KV cache.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs KIVI: a tuning- free asymmetric 2bit quantization for KV cache

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.751585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.941514Z digest=sha256:0095511d7a22adca0f780e3dc0af1e4ddd1a1fe7234b66495946387ff446762f

Observation 288d86ab-58a8-495f-a228-5b9ee5ec114d · outbound

This paper cites LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:06:38.123954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.947046Z digest=sha256:e99657bed855e263c5b91a3a80895d0aa94ab352940582f062b3fad49b954174

Observation db751f23-2830-4e43-b0c2-665b4dc94514 · outbound

This paper cites Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.951435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.951435Z digest=sha256:5417a60bb93114a3873caad63bbd05fecf0fdbc68e02729eb2fb5a860e73b8a6

Observation e5083c6d-33ab-49cf-825d-b81d20220c1c · outbound

This paper cites Transformers are Multi-State RNNs.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Transformers are Multi-State RNNs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.609049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.954862Z digest=sha256:7203d54a56e1f74c76660f4496d418ec1e98679e8e28e3d7f3be901dea7774d9

Observation cf1c9358-62f9-4a64-bdc7-b1ac7d5880f6 · outbound

This paper cites Scikit-learn: Machine learning in Python.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Scikit-learn: Machine learning in Python

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.500139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.958718Z digest=sha256:edbb79150e470ff8c0669c10dca136d121f31c815e03929403c60dbbc09280a3

Observation 93f5acea-eed0-44a1-a7f2-d3b38187eb1b · outbound

This paper cites Web-scale k-means clustering.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Web-scale k-means clustering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.418351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.961351Z digest=sha256:ddee1a09f9b1c91d62684889a4c8bd629f3852deeffa87249432dc0c2b8c1a99

Observation acf8426f-15ab-4ec1-bc9f-acf90e0f47f7 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Adafactor: Adaptive learning rates with sublinear memory cost

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.338961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.963935Z digest=sha256:5f7bc1a1c979277ee08b26014d713dcad75e875a4cd80aa1f7530bbe4452294c

Observation d8d2ff7a-a944-4ac6-88d1-a65a44d82955 · outbound

This paper cites QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.207679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.966289Z digest=sha256:263b71b1906dd04046e4c9860e636e60e34c1af68d84970f86efb3d674b1dea9

Observation 9b79bf44-ecb7-4a4f-b923-8ee1d1742da8 · outbound

This paper cites AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer- Wise Asymmetric Quantization Configurations.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer- Wise Asymmetric Quantization Configurations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:39.135861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.970390Z digest=sha256:e44191d8f3cf15a890259b6c594b9954617bb61d4124b57d87cb9b6bc5db51d2

Observation 4a804256-821c-4af7-941a-011ea97cf392 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.974660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.974660Z digest=sha256:228561c59496ec683b6252f43db03113545f50a78af3d60d943910f342b580e9

Observation 4b2a5842-f14a-4cf0-827c-9eb9dfbe3fb6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.978694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.978694Z digest=sha256:e333890f095c8f9f6fbc2f2db2d79e063aba5cfea9151bc0862df3b5ef37f6b0

Observation 11c6413d-c09a-4c77-9d51-63c2b6c1ab0a · outbound

This paper cites Attention is all you need.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Attention is all you need

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.904377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.981698Z digest=sha256:da7bdeda3135321136acb52077a73583e80b908dc2f7290d886c820f5576e6a8

Observation 98f37da9-ea15-4352-a0a8-bdd8e46ac2f7 · outbound

This paper cites Fast transformers with clustered attention.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Fast transformers with clustered attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.751720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.984746Z digest=sha256:c752afbc3d5ed66cf1d613b1f3bb49056ad6d8b03c0021d40cc71c4b2c7dd421

Observation 0f0aea8d-7a86-400a-859c-ac0a51eaf6ae · outbound

This paper cites SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.630982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.988527Z digest=sha256:0d60a9881c8042896075d4c4e0483bd181a785ae9946861fd56d2b9c0b5ba2c8

Observation 1c1e24bb-8c80-462b-a0a0-6e64696e4b37 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.990797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.990797Z digest=sha256:2e420c3ba7a2ee30cbdd393381a2d2e68266e22ac942e964cea62250b2411395

Observation a73e0ec4-c5a0-47a7-8d9a-0825d23228ff · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Efficient Streaming Language Models with Attention Sinks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.560300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:37.993929Z digest=sha256:0cde80b2a36c780e7127b17116cce448c10479cf3d0378df03994bdec4b3acb8

Observation 58d002b9-2bf8-478d-9a61-9193fcab74e8 · outbound

This paper cites vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.998948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.998948Z digest=sha256:e5dc831c978cd0a85db5167405bda691e69663f117299b3de50e17567691cf8c

Observation 5e20b226-1612-4e36-adc7-a2a165d53f6b · outbound

This paper cites ThinK: Thinner Key Cache by Query-Driven Pruning.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.003459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.003459Z digest=sha256:b3aae042db4e561823ec3cd0655654537ef81602f75aac58299c245d89f83dbe

Observation c31314c2-8d44-479e-893e-6badf21f3034 · outbound

This paper cites A Survey on Multi-Turn Interaction Capabilities of Large Language Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs A Survey on Multi-Turn Interaction Capabilities of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.007767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.007767Z digest=sha256:4ac6a6a109a3b1ab7f502c70902f9db24e3490a573007040883be80007f88aaa

Observation 4fe2a94b-3a14-4169-a780-0272201307f5 · outbound

This paper cites Long Context Compression with Activation Beacon.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Long Context Compression with Activation Beacon

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.011516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.011516Z digest=sha256:959e878fbc049ed0c890c1d393a678f8d1cfe3bd30d47677dbaf70fa70ddd162

Observation f2ae781a-3c43-4df4-9adf-74184f70c325 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs OPT: Open Pre-trained Transformer Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.014118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.014118Z digest=sha256:c11ea83cc1d987f69327566dd59a4640ef2fff71baa5312b995d6e4b2fa0eaaa

Observation 31fd5540-f7d2-41fc-9bd2-b3a1764fdc57 · outbound

This paper cites KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.481242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:38.017818Z digest=sha256:1be256898475ccc2e32a8a07479a62a5a2632dd2356fe61e19661e7e1a5b66db

Observation 7cb6c8ec-a7bd-4690-8ba2-937a9b53679a · outbound

This paper cites Chain of agents: Large language models collaborating on long-context tasks.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Chain of agents: Large language models collaborating on long-context tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.392823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:38.020684Z digest=sha256:21345a90b3cfa0fb64b96fc416c696780065bd01a10e0287384ca1430f574762

Observation 03539e17-272c-4263-9570-932b78338fe3 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:06:38.330399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T14:06:38.022879Z digest=sha256:676f398d29d6e626a5ec2121a77940a873fb6cef67579bd10b4d85cc2cea0a5a

Pith citing papers

No inbound Pith citation observations are available.