Pith. sign in

Paper Citation Record · LEDGER

ThinK: Thinner Key Cache by Query-Driven Pruning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2407.21018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.21018 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:22.978471Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2972a50a-33da-4f1c-a58f-be2f92ce2e4c · inbound

CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration cites this paper.

CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.978471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.978471Z digest=sha256:9b81df1e9a6c64829785d4c4e247fcb7cc971c3c62a497ae362687d3c62778c4

Observation f7b6bb84-7f31-4437-9796-8233bc028077 · inbound

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration cites this paper.

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:43:26.032859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:43:26.032859Z digest=sha256:31b75f91c07e48268cc61c2d28c9a9ba23b07860cdfcf6ebb4e3d247782fb554

Observation 4ba0e51a-3a24-43a1-a20c-9cccf4b082c0 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:09.833693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:09.833693Z digest=sha256:3216a53959bfc1ec6129f3e32b99f4093734d03f2abe6ab1f095161bb3e67e30

Observation ef238234-d80b-46c7-922d-49b515f2b452 · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.208166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.208166Z digest=sha256:9efe724c63a8ccd65aa9a9d0259a2276b239e2cb94439b2540795272ea11c6d9

Observation 5e20b226-1612-4e36-adc7-a2a165d53f6b · inbound

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs cites this paper.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:38.003459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:38.003459Z digest=sha256:f2d4a80a56cec19bdf974a74ec1a829aa5fbd652d2ba0b89932610e974808aea

Observation e854124d-b223-4855-ba5a-d24433b21779 · inbound

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding cites this paper.

StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:46:46.903274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:46:46.903274Z digest=sha256:db1690684d3ccaa37c24a69621e117ec7f2861b9e76afec800398600b2cc73d1

Observation dba18e8f-4100-4580-8983-76e9e9f9be50 · inbound

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection cites this paper.

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T05:09:31.978370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:09:31.978370Z digest=sha256:418a9b07245798640715cdecdb0a0d43e3e318bee366a8ce5641cf6904d36770

Observation 8d708051-880b-471c-a697-9f0ae88a4389 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.966728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:5590a00a5a99128fc982cd747588f689fc301236142808230e415edf56e83bc9

Observation 26fb19ad-10dd-4f93-9f83-babf9f55ba63 · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:16:54.241136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:180d3855bf9ea164d82d3d6dacdd0746e1919502a3b935557846b70727cb4649

Observation 5723fde2-d878-49d7-b681-b74db3eeff74 · inbound

Graph-Guided Adaptive Channel Elimination for KV Cache Compression cites this paper.

Graph-Guided Adaptive Channel Elimination for KV Cache Compression ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.685712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:06:08.786535Z digest=sha256:be7d26eb701f82ea98f7be6c6f78efc0d0d049ebeda46b2431973b4e5e958f4f

Observation 1160015b-699a-4e9a-920e-ed55aeb93634 · inbound

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache cites this paper.

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:33.256917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:16:15.869198Z digest=sha256:1d2f3d6492aa8cc85d851ff596aee39b302bc579bc4ecd779024fa6a3407635d

Observation 9134029d-0480-456d-8647-d70cd16db31e · inbound

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference cites this paper.

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:43:23.782224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:43:18.828740Z digest=sha256:180f286aae8754e6816202780afb874479cdc9d02124144ad616e612baab600e

Observation 7fc0bda9-428a-4430-b68c-5fab57119343 · inbound

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression cites this paper.

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:44:36.817269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T10:27:30.344430Z digest=sha256:ea6eb1a9643b4552154a78237e948bef49a8bceb562607dff9bfb165c73a0740

Observation eacde308-75ec-4038-807d-3d39bb8abc6e · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:33.643581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:33.643581Z digest=sha256:0014752c3cdd779df2b38f28e789eea7b04b85dae810552116afb40db24b7d50

Observation 02c7c400-91fa-413d-bb5d-e586ddb51650 · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:52.942535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:52.942535Z digest=sha256:819c468526c62030c9c667e8363dd9c134fd19d9c65193911da13213ef9cf9d5