Pith. sign in

Paper Citation Record · LEDGER

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

As of 13 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2607.14431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14431 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T02:10:50.027276Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcc0f714-5a6a-48f1-8a67-3630cf0085b9 · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.534551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.534551Z digest=sha256:259c50e4226f34cf1423af919c956b3df685c9fdd921566cbd80cf1cf9b75e3e

Observation c9ca5997-fc7b-41b6-8a17-a805ad298c9b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.822456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.822456Z digest=sha256:a4be0f7cbb342470ae206bd0e379b6d8f0f5641fc7962e7b00877960142429d7

Observation e25a3900-d4c9-4931-848e-d733724b97b2 · outbound

This paper cites CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.918894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.918894Z digest=sha256:9314f07b8ca82b346ad7acbe7eb8db4efcf7aff8164993ed532cb53e8fce68be

Observation 07fac811-4fb5-4a0f-974c-5002f3296a5c · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel SGLang: Efficient Execution of Structured Language Model Programs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:50.027276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:50.027276Z digest=sha256:97160531be213c71ce1ad1d05dfec9ccd97309a3b0effedb13cc868fdf10cf35

Observation 2fd66561-b457-4214-926e-88c6a9ab89a0 · outbound

This paper cites Energy and Policy Considerations for Deep Learning in NLP.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Energy and Policy Considerations for Deep Learning in NLP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.709223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.709223Z digest=sha256:14b8a0d7e54b0fd6c87bedc820cb036b5f0749d142ca02bc27b25ab964b349c5

Observation 156c3498-366a-4eb3-8ae2-12a5cb9308ce · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.471962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.471962Z digest=sha256:4a4af69243c745909e372cb66415a027aeb5420918d97159f055d36fde4744da

Observation a13ebabe-ba70-4929-aba6-63c486ffc31a · outbound

This paper cites Prompt Cache: Modular Attention Reuse for Low-Latency Inference.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.354541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.354541Z digest=sha256:8a96c4e0bae5818ce69413d041b86058c0604fdb0f01f06b3fc145213b124acf

Observation 0d19c8ef-1365-4605-88ae-b3ed6e986b9f · outbound

This paper cites Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.607668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.607668Z digest=sha256:3dd231eb24602c3353701e6c3756c10476f907b217715d30aaa280a859491f4d

Observation cfbb21ea-747d-40fe-ae00-09e9c6d68dde · outbound

This paper cites Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.265127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.265127Z digest=sha256:1f1e8e88e4b5e191c752ad9869281183ef831773b53f8f98ead1a8c40542ddad

Pith citing papers

No inbound Pith citation observations are available.