Pith. sign in

Paper Citation Record · LEDGER

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2605.24786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24786 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T11:52:50.864071Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T10:41:08.967261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1bc029c-2c26-4abd-9883-38f9899bba81 · outbound

This paper cites Longformer: The Long-Document Transformer.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Longformer: The Long-Document Transformer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:54:38.180010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:e76d9de0a48fdbc75d91cda192d7e7e6971d0df17533ed8acf57a2d26be7880c

Observation c17f3790-9252-48b8-8446-e0d1ec4cd2a8 · outbound

This paper cites Don’t waste bits! adaptive KV-cache quantization for lightweight on-device LLMs.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Don’t waste bits! adaptive KV-cache quantization for lightweight on-device LLMs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.910956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:8ed677540552ae736fa90dd2a252ded67b33b7a52a89afcfd7d2d4f439d4d9b8

Observation 936ac0db-5d5c-4b52-8df5-32a70327af2a · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Lee, Deming Chen, and Tri Dao

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.921912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:7f5ba6ce201d48f552265bfb758f3d726e9f7fa85ec2f500846bc2870037bea3

Observation 6cb3f8bb-9b81-41a5-a48b-b3de21a55b6d · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.177115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:d4cf5bb8a845d9754d8df47e0ef947db18e121700879b41dec29f4887c9b6bf3

Observation dc8d1351-2e37-46a3-8edb-8f73e70e7303 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.929325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:85e7b46bf5715ba7311ea377e008be3d1c64feb52905350cf8e3cbb12c2cc788

Observation 641b5c0f-e1d9-4a60-a41d-e9d1e67cc7c4 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.931052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:094ac8d11e13145adf1819635dd971cf10a1315973b95a5e630a2702ce7efb91

Observation 909b1b3f-23ba-48f0-a06e-d52acd190257 · outbound

This paper cites Kevin Zhou.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Kevin Zhou

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.934358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:4d2d92c1081c7d945773aceb92ae9afbe6c3aadfc1b675da80a2ee8b6b420f88

Observation f107b549-213c-433a-b995-53616b3c2761 · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for LLMs.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Model tells you what to discard: Adaptive KV cache compression for LLMs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.927548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:94255879c6b78df3a3101ebf8a06d137d321a4ce3ba0b3c764e17647b7f2bb34

Observation 0b79e633-e71b-4028-9780-cbe1cee68b58 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.939878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:148b6c822a8872f33b7eabb8fa202822f8ebfd8165ceb04f393073bfbb5cfca6

Observation e1d9e33d-82e6-4438-bee2-0f3608a74bc1 · outbound

This paper cites Needle in a haystack — pressure testing LLMs.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Needle in a haystack — pressure testing LLMs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.925707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:4aa3d4442c6125d41f630df463ac7f15af53c1cd02f04ee3b5f56595f696a283

Observation 237a378f-5692-49ea-b0ed-094b92a47f0b · outbound

This paper cites VisualWebArena: Evaluating multimodal agents on realistic visual web tasks.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.923919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:ac0baf5ed1bc9f55a0e7bf3044ca8eae2aebd4a1932404d5ac4ff3a40841c51f

Observation 0af25a82-f27e-49c6-90f6-edb0fcdb56df · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Gonzalez, Hao Zhang, and Ion Stoica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.935985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:d9210b565c840bf05a0c7d4b327136937d3658f715f1c1ee14af424620f3bd59

Observation b7a13331-c404-410f-9386-a80dbbc0cbe8 · outbound

This paper cites SnapKV: LLM knows what you are looking for before generation.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM SnapKV: LLM knows what you are looking for before generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.914598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:ab9b1b3b315b8422c496f0cfdd598060eed498f73a8614300ec64de00c0f7e50

Observation c0ce2a79-b786-4449-8e17-c83d845f1f43 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.916347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:a991765550c3f09d9221bcd024e8fdb98f2d9f0abd3e18a1892277255fe644ad

Observation b38c4277-fda9-422d-9f2a-cf8726c25072 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.918376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:1de8c1d569d8e8d2ff192a30d86c50391b66d415ee33d214697457bc9ec689fb

Observation 7d12ed73-7d97-4316-8c85-bfd7b3428f0a · outbound

This paper cites Pointer sentinel mixture models.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Pointer sentinel mixture models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.938125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:b7f66ca1000ad7eb350b7f5f70e8d3b91e7e43a54b0682358e883ea0aeda9798

Observation 3490ae96-86a2-406a-9f75-e829982ce10a · outbound

This paper cites SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.908972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:0b73d0351249d75cd14ec2890d663f1959e3fd1bab889c712520cf6b0799753e

Observation 120e942b-0291-4d91-a5bf-737daeba8ec0 · outbound

This paper cites gpt-oss-120b and gpt-oss-20b model card.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM gpt-oss-120b and gpt-oss-20b model card

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.950715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:e64889a7409986674218dc9029cb96fe32a41c77b599a9d2966f36ef1021cea9

Observation ac59b018-36da-4654-bf43-a9abe4f800a4 · outbound

This paper cites Language models are unsupervised multitask learners.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Language models are unsupervised multitask learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.912724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:d64a3d828e6b2193f47ec508488987781deb0139d19a8524ae4836be09c73cd3

Observation 1b2c2939-76c0-4124-9fea-5f7d70d2c013 · outbound

This paper cites SparQ attention: Bandwidth-efficient LLM inference.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM SparQ attention: Bandwidth-efficient LLM inference

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.920264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:cf724ac95f50e10a71b3134686abababa3baa338fffa353393cc242d7dcce533

Observation 51385273-83b5-4d4b-bd22-7ab2a4d072fb · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Tran, Yi Tay, and Donald Metzler

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.932662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:f484b1c7ad649e8c797d619e496012f7d167dad37370e89a3a49af076dd95984

Observation e5bb450e-3763-40a6-ae8b-aa86de536add · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context LLM inference.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Quest: Query-aware sparsity for efficient long-context LLM inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.947048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:578b3cd541997c41b12a06d0dd5ef732b706ac9c1682821d6a003e098f1f1440

Observation 26294ba7-a41b-41ff-bbb9-da6889a95bbb · outbound

This paper cites Efficient streaming language models with attention sinks.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Efficient streaming language models with attention sinks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.948943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:87b656d3644fb324fbe96e93b5dfe7e6f8030b4b24ee19a4fd38f758bdd82b8e

Observation 290d5321-04b0-417b-a409-aed13e10db3a · outbound

This paper cites UNComp: Can matrix entropy uncover sparsity? a compressor design from an uncertainty-aware perspective.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM UNComp: Can matrix entropy uncover sparsity? a compressor design from an uncertainty-aware perspective

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.944997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:614bc72b22d7d0298de6bb597b8da57dc3bbf287c01d7116a7aa2dbc6965b279

Observation 6eb7760b-962b-459e-92b9-0945d84f751d · outbound

This paper cites Qwen Technical Report.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Qwen Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.184689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:603e992e54dadb70d829a1d0c6cf9247a41cadda7b423538eeaa47a15dcaabb4

Observation 54c63ac9-97eb-4811-aa6a-22b6294eff83 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.941651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:6a2ca8422c48c724067cf2f4a0589b9fda81fccc328cef55530469c3c3ccb4f1

Observation fd1b4ba8-56e7-44b9-ad00-d1000099beb8 · outbound

This paper cites ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty

Reference 27

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T11:54:38.182377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:ba31f5769b7587e0b36ef29488f5dd5545ec4ae0fc20b9c83c7b0dbbdb238f32

Observation 4db70534-4856-471a-9c90-e36e08fb79f2 · outbound

This paper cites Justification: The work does not involve human subjects, so IRB approval is not applicable.

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Justification: The work does not involve human subjects, so IRB approval is not applicable

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:06:04.943280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:52:50.864071Z digest=sha256:e7baae9d9eb3476963dfc1d258dffe33f5c6ff031ac233a7bfbad6bf17cede3a

Pith citing papers

Observation 9b41a0b9-8670-4a03-8d63-59bc1651e080 · inbound

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference cites this paper.

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T10:41:08.967261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:41:08.967261Z digest=sha256:7186d3eda8ab47163f49dd8ee738148188b05f5822a46acb9dce6248d465813f