Pith. sign in

Paper Citation Record · LEDGER

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2605.24786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24786 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T11:52:50.864071Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T10:41:08.967261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1bc029c-2c26-4abd-9883-38f9899bba81 · outbound

This paper cites Longformer: The Long-Document Transformer.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Longformer: The Long-Document Transformer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:54:38.180010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:aa55c97f80b8fa494a4dd7e6a251974fab893a68d6c70486bd6e41bbb7887989

Observation c17f3790-9252-48b8-8446-e0d1ec4cd2a8 · outbound

This paper cites Don’t waste bits! adaptive KV-cache quantization for lightweight on-device LLMs.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Don’t waste bits! adaptive KV-cache quantization for lightweight on-device LLMs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.910956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:41fd6e2b0555f4d66943737fe5915aaa5450dbab14195941127d9617ed7baa2d

Observation 936ac0db-5d5c-4b52-8df5-32a70327af2a · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Lee, Deming Chen, and Tri Dao

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.921912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:b34e5e9484a82b93790b39ac5a2a7065fb7c24120aa4dc1adc5ff061757c3812

Observation 6cb3f8bb-9b81-41a5-a48b-b3de21a55b6d · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.177115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:9f51220c23fa7645057691ca289553a6f70dc91239a03c5ceeb0bba20a7b3e2e

Observation dc8d1351-2e37-46a3-8edb-8f73e70e7303 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.929325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:58c0b9eff27ea63a8311ab1881a68d3e08097dcec7b0f0dd74c91a9d2a0fcbee

Observation 641b5c0f-e1d9-4a60-a41d-e9d1e67cc7c4 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.931052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:66934b41c404375e998876a4d19b2b813a5c53265369c14594b9f4387b18c33f

Observation 909b1b3f-23ba-48f0-a06e-d52acd190257 · outbound

This paper cites Kevin Zhou.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Kevin Zhou

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.934358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:9f434e64563979d40a52263212eeddf1032634a2e04c4d33d0aa76b93fe26271

Observation f107b549-213c-433a-b995-53616b3c2761 · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for LLMs.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Model tells you what to discard: Adaptive KV cache compression for LLMs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.927548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:7acdbd9974a87dcdfd3a29b89a96a8e9c104a20319c42f7da34825d1e319677e

Observation 0b79e633-e71b-4028-9780-cbe1cee68b58 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.939878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:5301daabbadf92526bf02c02d9936f107d05771236af6249f63891fd70f952c6

Observation e1d9e33d-82e6-4438-bee2-0f3608a74bc1 · outbound

This paper cites Needle in a haystack — pressure testing LLMs.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Needle in a haystack — pressure testing LLMs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.925707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:6ea7c15f1a1ea57d4d609bc16fde0d1157ac776a937d79fcb8cf5eb830ef37ee

Observation 237a378f-5692-49ea-b0ed-094b92a47f0b · outbound

This paper cites VisualWebArena: Evaluating multimodal agents on realistic visual web tasks.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.923919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:f777e2e30581f6089d4f7e6626691cb93162cef4c3850c0d8900cec16b01eac3

Observation 0af25a82-f27e-49c6-90f6-edb0fcdb56df · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Gonzalez, Hao Zhang, and Ion Stoica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.935985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:5d11e1a8b19153e231811021980d3ba4f27d38f412d7636277efe48ac47f7e28

Observation b7a13331-c404-410f-9386-a80dbbc0cbe8 · outbound

This paper cites SnapKV: LLM knows what you are looking for before generation.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM SnapKV: LLM knows what you are looking for before generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.914598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:bb849641ed49f385677a5b84c630c6e558050e5938227b1616154f09df878403

Observation c0ce2a79-b786-4449-8e17-c83d845f1f43 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.916347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:052431514239b7fa28c6060a94ea2b5a7e228d3835e523b5dcda6e0112bedfe1

Observation b38c4277-fda9-422d-9f2a-cf8726c25072 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.918376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:9b8721d8016f3d6e8ce37de7a5e2886700c955c2172598a5bae623e1f96e7b49

Observation 7d12ed73-7d97-4316-8c85-bfd7b3428f0a · outbound

This paper cites Pointer sentinel mixture models.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Pointer sentinel mixture models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.938125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:def9090ac222b8f4b90810395ecc433e33a662d813347c57e407210cb1085500

Observation 3490ae96-86a2-406a-9f75-e829982ce10a · outbound

This paper cites SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.908972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:87d56e22b1a56328db16e2efe131e2ed4a6999ca2ef13726c9c4c0649e202d03

Observation 120e942b-0291-4d91-a5bf-737daeba8ec0 · outbound

This paper cites gpt-oss-120b and gpt-oss-20b model card.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM gpt-oss-120b and gpt-oss-20b model card

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.950715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:b0a247c26d8badbf65da5bd63ffc895f03c12f81583bd9fb389041f51eb798a0

Observation ac59b018-36da-4654-bf43-a9abe4f800a4 · outbound

This paper cites Language models are unsupervised multitask learners.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Language models are unsupervised multitask learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.912724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:54207f69c14adc80fa759e16351fba3d2415b178001e5c5c9523e6fae50079ba

Observation 1b2c2939-76c0-4124-9fea-5f7d70d2c013 · outbound

This paper cites SparQ attention: Bandwidth-efficient LLM inference.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM SparQ attention: Bandwidth-efficient LLM inference

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.920264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:2b137b76dcc1a83601f42aa43ac6a2bb48442683c0cdc4ab214cd31426fa1732

Observation 51385273-83b5-4d4b-bd22-7ab2a4d072fb · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Tran, Yi Tay, and Donald Metzler

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.932662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:03073240c9270580afdb8e70b687e0c0b083ebd937c82e1cb0a92b9910570c06

Observation e5bb450e-3763-40a6-ae8b-aa86de536add · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context LLM inference.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Quest: Query-aware sparsity for efficient long-context LLM inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.947048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:bd40f919818bb72ddca482ec015b59b1ffb5b2d207dbdd478dff85528141418a

Observation 26294ba7-a41b-41ff-bbb9-da6889a95bbb · outbound

This paper cites Efficient streaming language models with attention sinks.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Efficient streaming language models with attention sinks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.948943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:8a28da871a3abbc64eef09c4629bbcd4e36a25785e0422f15c33824d61f2e469

Observation 290d5321-04b0-417b-a409-aed13e10db3a · outbound

This paper cites UNComp: Can matrix entropy uncover sparsity? a compressor design from an uncertainty-aware perspective.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM UNComp: Can matrix entropy uncover sparsity? a compressor design from an uncertainty-aware perspective

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.944997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:8bf2f73078ae8213d8ee9231da42217c8b45da0f44af4ce69eaf0e81f15e6722

Observation 6eb7760b-962b-459e-92b9-0945d84f751d · outbound

This paper cites Qwen Technical Report.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Qwen Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.184689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:66d3a9cee170140512b40cf2ee8c958dc46935a7832d20b22dd17a282068134f

Observation 54c63ac9-97eb-4811-aa6a-22b6294eff83 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.941651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:bc198420587b79b579656701ba5b61edd5e08d986da37522317ff4710cc9df93

Observation fd1b4ba8-56e7-44b9-ad00-d1000099beb8 · outbound

This paper cites ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty

Reference 27

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T11:54:38.182377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:26272d7666b3d38e9eb624c014094c6d772e0eb7505a3248e337f7302ed8fe3c

Observation 4db70534-4856-471a-9c90-e36e08fb79f2 · outbound

This paper cites Justification: The work does not involve human subjects, so IRB approval is not applicable.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Justification: The work does not involve human subjects, so IRB approval is not applicable

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.943280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:68a1765af54fe839b45848572731eb1b80b33da6f16e60f29e664ee0837aab1c

Pith citing papers

Observation 9b41a0b9-8670-4a03-8d63-59bc1651e080 · inbound

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference cites this paper.

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T10:41:08.967261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:41:08.967261Z digest=sha256:7186d3eda8ab47163f49dd8ee738148188b05f5822a46acb9dce6248d465813f