Pith. sign in

Paper Citation Record · LEDGER

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

As of 16 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2411.18077.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18077 v3

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:37:07.697300Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:47.483457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T20:59:01.556666Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7ad4eaf-d1d5-4d47-a479-334cd7017ad5 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.114579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.544023Z digest=sha256:af24537cd5b4f36bbc1a3bce5d2c639a555efced118e73fea74ced5b40cb24d9

Observation 5af67f7f-5cea-4d68-bef3-4751fcbf1add · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.548481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.548481Z digest=sha256:4c5afc1762c41f4136c3ec629af02cebea082c346c17150354024565d56d122b

Observation 592c3a29-2acc-4656-848d-50800257f37c · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.553067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.553067Z digest=sha256:379e9616bdac247edd8741aa5275a74f68450b2227d480282bcadf9be6b178b6

Observation 7d456d26-f44c-45d2-870e-b89f54a6f0c5 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.557293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.557293Z digest=sha256:aa0d3c3c02116f3ed60103d2486aa296aca14a37adfc342ab3e6c275a56b0603

Observation 03199717-1dcd-425a-9376-503a74745195 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.561541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.561541Z digest=sha256:31128b5a0874ef515e1273f97aa1f75cca66335ca9bc3bfc6c3354a54ee0b646

Observation 199d81b7-0328-4935-ada3-998fede18f89 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.565472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.565472Z digest=sha256:04652ab4251b969b1dbd001ec6f320deba6b369eebc6866c1e7259e8c6efd415

Observation e8433fc1-ed3c-4e87-9e95-4dd9418229ce · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.570007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.570007Z digest=sha256:96052afe35fd109e380ae6b7151c8fc03ce53bab340f1714381b0d77eb3c663c

Observation 32394cf1-ad8c-4554-8e12-3cdc54a95116 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.573793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.573793Z digest=sha256:1b8595481b6786422256f529cdd1b282397af69b7a06354931b70b18ca1b905f

Observation 4cf2ff58-02f5-4c4c-91b4-2c10cc73130d · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.577744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.577744Z digest=sha256:9033cc7bfa7bd82d2478b8f49761e2454d272eb00937c5ceafcbc0e5a4daf15a

Observation c6ce3c48-7d7e-4ca0-840a-3a7eb8132d43 · outbound

This paper cites Mistral 7B.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.581390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.581390Z digest=sha256:72362a57853a9501785d9e05e53124a8eb4433a2d018e85b46233fe837902b81

Observation 95b0495f-cc05-4fd9-b454-a6d9643c4add · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.096814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.585359Z digest=sha256:2590b2800bcde3266db02e5c5d44917278b93d86f76bf5ecfab48e6a9c9b90b2

Observation 904d9b76-48d9-40e5-8e95-fd183426b773 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache SnapKV: LLM Knows What You are Looking for Before Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.588847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.588847Z digest=sha256:41800c38be07b091894d324c2451f0a7372da42368475d8896421a72e72b4d72

Observation adbad6c2-b9d1-468a-b410-838668088066 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.592449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.592449Z digest=sha256:b21cb66ba0b6ab8681f01e5032793bdbd80e679db2a042f138bfe09f8261c3ac

Observation 39ffc95a-dca4-42ef-829d-d6bc784359bd · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.595756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.595756Z digest=sha256:3a9d3d2453abe70d84f315f968890ecc413fae5f069e6152267964e7ba377a51

Observation 4f7676d9-5061-4d50-a2a4-ea539a3a263e · outbound

This paper cites Multi-head or Single-head? An Empirical Comparison for Transformer Training.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Multi-head or Single-head? An Empirical Comparison for Transformer Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.599724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.599724Z digest=sha256:f4f3a7d9c92044a9e69d18f75598cf3c7747cbbe519d5fe817a31ee387649ea5

Observation e1734be7-5d10-467c-9a4d-755149e86b5a · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.603254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.603254Z digest=sha256:a607d6bfd20654e25b84fdec71976ef19d75e897e9b41c892ae67c0037d293b3

Observation 7ae1c708-e961-4a3a-b887-930d580ce0d4 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.078454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.607108Z digest=sha256:b359f0d91004131f4957502e0d698ea742dc055a6b8129373bc583cf75a709c1

Observation 1f830ab9-1b00-4ec0-95da-5b50c75c7476 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.610637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.610637Z digest=sha256:eec2ae0f2e520be6817e5623fda8fa9fb7b8f26360e74f6add72761ce494a608

Observation e45af902-9bf3-44bc-bf9b-cd7f903b4e55 · outbound

This paper cites Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.614689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.614689Z digest=sha256:801cc123d109290888593c5104d114d060335556bfd22606337998b65624c05a

Observation a8b8127c-4618-4611-bc2b-9dbea5525753 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.067490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.618605Z digest=sha256:52bee9040d8bed56c7d1837dccd1c4d6ea54a33f5abd9e6fd71a22d695c23f16

Observation 7057606f-fac1-4ad3-aeb0-15f972b7e1ae · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.622246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.622246Z digest=sha256:e7a71686e81af1d812bfaffd435a1e390746c976958f9e5dc5b18285d561f622

Observation 57c805ba-a686-4df7-9460-27b12a7190c3 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.049427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.625766Z digest=sha256:1c9e3f17f9e7b138382d1a35c4dac2625f6ec18dd60f29c72374018d246dbcdc

Observation 3fa7a7c7-6b28-4d72-bb1f-86ea6ca9b1c5 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.629296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.629296Z digest=sha256:725fd7e0b2947cbc314cfcd1fa3e6e16fe9f20511973d8e9b2aeca700eb8f9af

Observation bd394d09-30f7-4f6c-bfe9-99f30fa42aff · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.037753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.633190Z digest=sha256:8622c361d346cc6356ae2034964395aaf3167e10c59fa8f12747bf0119d4f999

Observation 80ae3419-7cf7-48bd-860e-22fbb5e45893 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.636669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.636669Z digest=sha256:6de2edc227740ce417e31aeeb2ad9233689c79e0859859968b3fe90ef118f9a7

Observation 3788dc32-33fe-4100-9964-5c88c6e2a9c1 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.026080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.640473Z digest=sha256:faee3ee89d600474a7080184df47e367887918cdc4f8cf13bfe73e0ddb3c5054

Observation 97a6cc9b-743d-4e2d-8581-1a9877b43c1f · outbound

This paper cites Do Large Language Model Benchmarks Test Reliability?.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Do Large Language Model Benchmarks Test Reliability?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.644899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.644899Z digest=sha256:76a559c24a2b5d9b59514af71d9da70da8e146948250a90506baddc3f1f20595

Observation 7e18dec0-c176-4357-b125-3a597f77447c · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.014635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.648882Z digest=sha256:ae93843e59aa5119928fa6d04973b7750c42802f527a6d4285015f6a8e363d99

Observation 04249263-2fab-41bd-a950-dabe643a7754 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.003790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.652446Z digest=sha256:91e6d61be51a4460bde7e7dccebdd5b09c8d7608c936689d34cb8fdbe4048f8e

Observation c5ca8fff-152a-451d-aea9-2bebe4f1518c · outbound

This paper cites D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.655879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.655879Z digest=sha256:dfda30edce4bc4701809c6cb05a26d89966d09503d3899bddaad3765875b737b

Observation e6564002-5de1-46ec-94b7-1698ddbae20a · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.659393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.659393Z digest=sha256:2e3253420c6b295840b8fadec7212e77fc66b043984843828a7b2c00006b71f2

Observation cd0916af-31f0-4431-ac51-2fc67cc52b04 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:07.992453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.663004Z digest=sha256:f559635faf4342b45ba8b66763a5652270f3e8b735ffbdb5e654bbfa10c68cfe

Observation 20e63b59-4d5b-4e35-bafc-1cbcebf0873c · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.666624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.666624Z digest=sha256:cdf108ae022d03883e8a56c8db27f274c963c55d288994f0e66cbc6f3157e42d

Observation 6c22d63e-a693-4633-bfd8-eb0c8e03e13b · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Efficient Streaming Language Models with Attention Sinks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.670319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.670319Z digest=sha256:a4ccc2bf71600a1921fa312899b7a6d7c713553ed66403f7586248a40a31a6f9

Observation 7ced9945-c15e-4e04-939c-711b8dbb8902 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.674267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.674267Z digest=sha256:4e4ab8cec0512bf4cb147bd7c2b1c329c8109fa5255950d16eedde5c5059d9ca

Observation c9d88fdc-0c2f-4b57-87f3-7a4dfaadb303 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.678061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.678061Z digest=sha256:e37e2fbd9227eca9f50e5482a53ecd40dd256da2243237ba082bc9f712ee7f94

Observation 7267b29a-8ed7-424e-b943-25a38825d5d3 · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.681788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.681788Z digest=sha256:3b0cb6d1a2425358c62a4b074e19155cbf5e86ef81e9031796d924025add21f0

Observation 9340b39b-65a6-4abf-adff-3f0066468f2c · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.685940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.685940Z digest=sha256:2c6013d26e7b64e796ef4e586e173494695de229a87b893c3bfe95019abb5218

Observation 248e51a3-32a2-4a0c-b655-ab701a21bfa8 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Barrett, Zhangyang Wang, and Beidi Chen

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.689428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.689428Z digest=sha256:902481628b6566146b7af0beaca40789fe407ef13a86331597c395c94bcf4052

Observation 93f5e1f6-0042-405c-9339-19fa85f70aa3 · outbound

This paper cites online" 'onlinestring :=.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache online" 'onlinestring :=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.693059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.693059Z digest=sha256:5cf5f86c48baaede8f5c9053e38eb6ca6471479ee1a6ad56c36d1026791abecf

Observation b66d035f-d673-4e4c-92c0-6329e84a19c7 · outbound

This paper cites write newline.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache write newline

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.697300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.697300Z digest=sha256:d0c0c07936b9ec63543248ffb6faf9c24b76b94c47472a41ef637ab3d0aca5e9

Pith citing papers

Observation b1956160-20f3-4901-b2c6-5b51d48b213a · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.483457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.483457Z digest=sha256:af1e48944a09f42832cc17ad01e14a311b85ff93482b044e52dc3760f8d784be

Observation c752e517-becb-4267-8201-5ca35c10a045 · inbound

Minimal-Intervention KV Retention via Set-Conditioned Diversity cites this paper.

Minimal-Intervention KV Retention via Set-Conditioned Diversity MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:59:01.558729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T20:58:07.403912Z digest=sha256:f3c851e7072f1484bc93986ff07e729507007185daca0d16419a8bd34a93019e

Observation bbcf9321-28d6-400b-946f-c5dddb42a72e · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:18:16.562435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:53d6bc35da7d41eae2c1f25e759206b9a8ae3be187dc88188a7eea383d50bb64