Pith. sign in

Paper Citation Record · LEDGER

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference

As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2604.19769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19769 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T23:46:59.863921Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:07.252227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29ddb2a5-6661-4574-8292-f5ea4a6da519 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference LongBench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.438410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:9ae9592154362a91b0bba98b0859814ae95fc8957fc323eff36c32923de952b6

Observation 72188809-63a1-469f-ae2f-4cbd91f52452 · outbound

This paper cites an unresolved cited work.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-14T23:48:19.434347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:6bb472345b161cd718402793ca399e670bcc3754109480d21dcb0e0deec60c48

Observation 031e9997-33ae-4377-88c0-6fac96767e37 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.430208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:d0f9dcd4b3907dc5af09f608ca79e5bd4172f56642ed71ab566982a9d3e90256

Observation c4b586df-b9e7-405c-8e79-9edf7d377d8a · outbound

This paper cites Ruler: What’s the real context size of your long-context language models?.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Ruler: What’s the real context size of your long-context language models?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.425546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:428c4515951cb8957bc639fa493f9eb4f37f12482d395484004e4a02601c9325

Observation bb6d5432-39d0-445a-a660-77f9b0b38cf7 · outbound

This paper cites Qwen2.5-coder technical report.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Qwen2.5-coder technical report

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.498603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:22abea814b8196d203637f794326516cc2f5531b1de690885dc4ce9279cc35aa

Observation 665954c3-1d4a-4ad8-b0ec-0231bf90ee5f · outbound

This paper cites an unresolved cited work.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-14T23:48:19.460331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:53e62b854643965de321308560c41c47c0499162638899b6fa1c60355ee0346c

Observation bbfb79de-7fd1-4f14-a545-8bd5ab549c36 · outbound

This paper cites KVPR: Efficient LLM in- ference with I/O-aware KV cache partial recomputation.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference KVPR: Efficient LLM in- ference with I/O-aware KV cache partial recomputation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.464413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:0dbc705328d64139155c0e55185baa8a583ae6cd9714481186558cece42321a2

Observation 78f74fd8-0de9-47c0-9e61-78802fb8e95d · outbound

This paper cites [Jianget al., 2026 ] Bo Jiang, Taolue Yang, Youyuan Liu, Xubin He, Sheng Di, and Sian Jin.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference [Jianget al., 2026 ] Bo Jiang, Taolue Yang, Youyuan Liu, Xubin He, Sheng Di, and Sian Jin

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.522699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:53f47e9dc3be9b409ceb6b56f22c44f664fe520b4d773c61fb6b130a8566c38a

Observation 8aec22ed-6b29-4f7e-af93-2730afc50783 · outbound

This paper cites Reformer: The efficient transformer.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Reformer: The efficient transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.476854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:c1085a522c09c614b51bf1f04e1c7bdf16a3aebbb441204a34d9e48f7e5bc2aa

Observation 4a9bde02-e4e0-4b1b-8458-99c1e830c54e · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.491644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:928547b243c368fd75af542a1c5a670ed846478d450bcf9596d36f36f5536ccb

Observation 32f5ccc4-85cf-425f-86b8-afd6adfd0a48 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T23:48:19.056976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:084625ba61b297339b44a0cdb6db21114dfef9b4399e0687f60f5d3a0b34edfc

Observation 3fb4c6b4-fb6f-490e-addd-e96cb520112a · outbound

This paper cites Freekv: Boosting kv cache retrieval for efficient llm inference.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Freekv: Boosting kv cache retrieval for efficient llm inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.451759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:2bfd5f6016da329571c75bdd048ec43b03a5a3fddf43019fafcd275f8102609c

Observation 80c4a8ef-c933-43cd-91e3-e8a9cb5aeeeb · outbound

This paper cites MiniKV: Pushing the limits of 2-bit KV cache via compression and system co-design for efficient long context inference.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference MiniKV: Pushing the limits of 2-bit KV cache via compression and system co-design for efficient long context inference

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.502500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:885c0697960b84c02032b71d6ff5a30a132fbc87ee299511c0296a46047007e7

Observation 8a5ebbf6-a19d-47cb-8ff1-76c3f7f96e58 · outbound

This paper cites [Shenget al., 2023 ] Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, and Ce Zhang.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference [Shenget al., 2023 ] Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, and Ce Zhang

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.495342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:082de2722bc044cf9ef17fa6ba82004c9821bd9d29004de922b7f2deb1126b04

Observation 155289ca-e674-4f90-aa5b-72ed8b94d694 · outbound

This paper cites Shadowkv: Kv cache in shadows for high-throughput long-context llm inference.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Shadowkv: Kv cache in shadows for high-throughput long-context llm inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.481161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:dc05c6f81964153ef5512d9fbc115294222d30c0b3fd46bcdf70f9b55bb79e31

Observation 998898b5-b396-4997-9aee-6206a0d8b1f3 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Llama 2: Open foundation and fine-tuned chat models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.485461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:b8954d1abc1a327d07c1b9461a7f845501d000641167ee848bdf8a12ec330bcc

Observation 1525ca07-aedf-4f8e-b6e6-15973a7367fa · outbound

This paper cites Leave no document behind: Benchmarking long-context LLMs with extended multi-doc QA.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Leave no document behind: Benchmarking long-context LLMs with extended multi-doc QA

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.549878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:7c5dbbd851c2a23aa4de2c14ce059c40a1cd13e210a41d114e0a8b6c3792ab40

Observation 811fba38-13a6-4b3f-8364-a0fdeb32bd69 · outbound

This paper cites [Wanget al., 2025 ] Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Tianlong Chen, and Huanrui Yang.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference [Wanget al., 2025 ] Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Tianlong Chen, and Huanrui Yang

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.557757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:93d2d01682cab88fd761d7fb105ae4b8b994ab07561af2c00c945539e0c7992b

Observation 7ff87bee-b022-4eca-a3a7-7a6a9ef61a1a · outbound

This paper cites Efficient streaming language models with attention sinks.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Efficient streaming language models with attention sinks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.447436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:27e09ed60e6eab916529fa343019005e28014c9d3c04f80c41b06c76c0127f2d

Observation 311491fb-a1f3-4324-a233-f63d5b7a9d7c · outbound

This paper cites H 2o: Heavy-hitter oracle for efficient generative inference of large language models.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference H 2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.472189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:b1126e2147ef956ed8113b41a60492cbb41d2bc8764fa174085eb899433f8e46

Observation 8632b857-300a-4f2e-b393-00db69115fa9 · outbound

This paper cites an unresolved cited work.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-14T23:48:19.442909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:bf44878505aaf4971b0ebac4ca6004c14524e0d0f12ad706b255892742dc0774

Observation 9f456916-6588-4b93-bcd4-7ba6904fd3d4 · outbound

This paper cites ProcessBench: Identify- ing process errors in mathematical reasoning.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference ProcessBench: Identify- ing process errors in mathematical reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.456353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:a313a969540d4ce8d84f2604672809d962e8f2a98f4b70f8b84983044c02881e

Pith citing papers

Observation c758b92d-0a3e-4a3c-9ab5-6252a2e02f4a · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:07.252227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:07.252227Z digest=sha256:78360a0dcfdd531eb0810d2c35a4f89970c7c796cdd614ba7329811b28ff7feb