Pith. sign in

Paper Citation Record · LEDGER

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 6 inbound Pith citation observations for arXiv:2501.19392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19392 v4

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.382301Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T04:51:04.068354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:55:24.203465Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ad5d31d-d34d-4027-8f5b-c90498d5144b · outbound

This paper cites write newline.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.163378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.163378Z digest=sha256:f547054b08ae88abd3202326d8f9a6a38c896c29ce00b4799f081808478fc5c3

Observation 251e4d38-cb7c-48ca-8464-a9042a9ff396 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.168712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.168712Z digest=sha256:0482fdd43f0d3bf52bbe387f0c33d42191384f9e1097c7559fc86aaee4376784

Observation af7d1eb7-9eae-4398-a32d-2dc231825daa · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Understanding intermediate layers using linear classifier probes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.172410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.172410Z digest=sha256:0c17ea026d176453e0dc83b7cb5056f0e1a112de01b637ac20bfe4a0214c31fc

Observation 3d464322-1ec4-4d62-add5-38a920dc34e3 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.176784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.176784Z digest=sha256:80de8e772ae2968b906ae5a8d8b52f81bbe7dbebd6a171755d9f5234aa45e513

Observation 3ab4ce44-4aa3-45d0-8dbf-615a79164ca4 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.181193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.181193Z digest=sha256:3626ffd44fffc213de0a306c69058c906646ff79f7cf93e1cb6247c9af6fc36b

Observation 4fae5ebc-01a5-44ce-8995-db483bb3304d · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 6

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T20:34:02.075389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.185806Z digest=sha256:20f59ecbfd32ef55c886b6ac2b1533ade3adc3befde3607049ea10a856c3f76a

Observation 9e72fe70-0f02-408e-ad53-19ad07a7a089 · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.189891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.189891Z digest=sha256:03d2017702d8e7905f098c3a408f96c36458ed2a25795ea8f6749c30a3a56e13

Observation db045350-6224-45c4-a0ee-2771b8660ea2 · outbound

This paper cites PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.194069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.194069Z digest=sha256:604bd47421d74ecd07b4a782c60e9b9a7335fee61411bf160577a03f3f1ab21a

Observation ed214d9c-c945-4b69-b3b3-c80e9c68c2e9 · outbound

This paper cites Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.197891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.197891Z digest=sha256:2565e4e049813939ec127f75d4ffe737eabd0392707f3b2d67b179cc2b45f3c4

Observation 7230ed0d-e4f3-4763-8368-4242e1735827 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.201633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.201633Z digest=sha256:86e859c7be54deccca5e6bef80e6d1b78e9568b8c17444eafe73c45f8d7fd68d

Observation 9ed2fe7f-cb97-493d-871f-444192003f39 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.205850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.205850Z digest=sha256:229dc59287d9baa95f70f824e0dd46b2d294b51fbe37408ecf30a66c264deede

Observation 4a74aea1-516b-4ebb-8de6-e9e59c0f7091 · outbound

This paper cites The Llama 3 Herd of Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.209552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.209552Z digest=sha256:b3f42d9951b0046ee61ed6c71610cfe4251df7f7c0283fb44fe5db796d3db681

Observation d792f3f2-6b06-45a3-a052-5d75ce161ae3 · outbound

This paper cites Towards Measuring the Representation of Subjective Global Opinions in Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Towards Measuring the Representation of Subjective Global Opinions in Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.212965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.212965Z digest=sha256:1b72818e4c32f85e8d05eeb289d3b861cac97571a89535e7751d881405a5d895

Observation 43925cb0-b1e0-4fa5-9897-df956ab7e938 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Extreme Compression of Large Language Models via Additive Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.216140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.216140Z digest=sha256:0e49ebc2df8e9957de2b177c272332b959150d48c9d0a8dc4fa76fa2f85f851d

Observation 8a374d1a-a87f-403a-b8ae-0f658c2b9086 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.220117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.220117Z digest=sha256:9b8f2b06987361ed05a57fdebaea334f64ff36f9d4b72e09bacdc635db58e529

Observation 8c7041fe-7771-41b7-8d3c-ab6a53ac0f1e · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.224437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.224437Z digest=sha256:6a69a08f6acbe7cbfe0d80febb2dafaa0d5e31a6d086d1a75ab828af0f9f9fae

Observation a2276e2c-5d4c-46cc-81dd-afd48dae68ad · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.220412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.228274Z digest=sha256:a4d342883f932d2bfd80d17fbeae963df5527b6bda777f005eb1a0c41a850fd0

Observation 95a1ce10-7f8c-4dca-8d26-f4291b42853f · outbound

This paper cites Fast matrix multiplications for lookup table-quantized llms.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Fast matrix multiplications for lookup table-quantized llms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:02.211651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.231283Z digest=sha256:db50c76129c2f06edfec4dccedd2d49971cc1bc024da1ee57b552d62528b29eb

Observation b8b6f30e-69ab-4186-804d-ceb69798c826 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.234695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.234695Z digest=sha256:7f488e9ca26d57e06ad6b883a678497d14662c2d5401ebe173797ad694e682bc

Observation 63704b6d-d0da-4d2f-9966-e36b374e5998 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.238595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.238595Z digest=sha256:695d4321fdb159b9a1728d21f7a2c4be9d391d878c48103bf37b0cfccbfd3b6f

Observation 82fe702c-a446-44c3-8773-9407f9878293 · outbound

This paper cites Stochastic distributed learning with gradient quantization and double-variance reduction.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Stochastic distributed learning with gradient quantization and double-variance reduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.242633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.242633Z digest=sha256:a8cd0f2daadf68809d8f4d2b7692cb188e305c66d5eed9361bea929cd155c1c3

Observation 4fa84ad4-54f3-4f07-8dc1-78f7ea39e1a7 · outbound

This paper cites Optimum-quanto: A pytorch quantization backend for optimum.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Optimum-quanto: A pytorch quantization backend for optimum

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:02.196286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.246441Z digest=sha256:f8e6730bc3e4c8ac0cb8bcbdfbaa07a3a5e7f9ff0cf971ffbea166e7783be63c

Observation d80ec495-591c-4b2b-98d8-d3be9eca364f · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.250786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.250786Z digest=sha256:7dd39753672532bf11a51290c11990f38500fbd23119f74ca7d4fb4751262884

Observation ddf7a799-c6fc-4f2c-95bb-f12f480a8f7f · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.254750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.254750Z digest=sha256:04a4561dbbb7e2536fed3220c650f9dc43b3341cf405b3826a4fd61ae711cfa3

Observation 363d5d4e-f8ef-4346-a9f0-3e4ee7b35342 · outbound

This paper cites A Survey on Large Language Model Acceleration based on KV Cache Management.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models A Survey on Large Language Model Acceleration based on KV Cache Management

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.257427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.257427Z digest=sha256:e044582adc4bda00b856f90474cf9e400c76d365b24559b0cc6019685964effe

Observation 3bf02429-cf6d-4568-9e3c-7552eecc6d51 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SnapKV: LLM Knows What You are Looking for Before Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.261682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.261682Z digest=sha256:0b2e1f556a84f6663490072816c589b468e11c7c71815a28daa9462b9f50662a

Observation ef36e9a9-22d8-4947-a948-1d7fb234a758 · outbound

This paper cites SCBench: A KV Cache-Centric Analysis of Long-Context Methods.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.265626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.265626Z digest=sha256:a52e2da37efe30bcc70ddaf36cb5d02af344f1e0543530171ffcfa0a9c235185

Observation f4d0dc2c-42d4-46ac-9b48-cf2b1d6fb4cd · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.269681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.269681Z digest=sha256:e3c1fbbf912deea55ed7cd42fb6d50d7076f983722aab588a14b6c1b59c68f2e

Observation 8100fd37-f4e5-42b6-8a23-f126bcce6ec9 · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.274336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.274336Z digest=sha256:2611c2c5d27457a19be03e2ad45c0c098eacf8699249531001892b72d4f8f61d

Observation a3c491b0-3ab9-45ff-a0f2-11633560713e · outbound

This paper cites IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.278582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.278582Z digest=sha256:1001062be47d1debd5b257489a702f1e18e8bbe2cb45fa1095ce03e9be5a4f3d

Observation 70c22dd8-2852-4c04-93f9-ee6de6058fd1 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.281832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.281832Z digest=sha256:f3b4659a14bc3d3b3ddeb038204849e160eee7ca3379f1c8010d52b110b29953

Observation be5b2b39-17d7-4e12-baa0-d2ac0ad6d211 · outbound

This paper cites PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.285840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.285840Z digest=sha256:580df6d705db96e2f9d8e315c65a611ba55f5ca2f5fc15edcb2d46dbbf2077f3

Observation f0f3d9b0-5619-4ccf-a814-6d412e9edcfc · outbound

This paper cites Pushing the Limits of Large Language Model Quantization via the Linearity Theorem.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.289119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.289119Z digest=sha256:fb16161b732bb8dfab8818e5fb3119a57fb850586099130a40ba685432da85fb

Observation ec4bf091-6217-4858-9150-f8330403e6bc · outbound

This paper cites Pointer Sentinel Mixture Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pointer Sentinel Mixture Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.293602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.293602Z digest=sha256:341290153b1f73b533bcc5008f236cb543708ab9f7fc5f8bed54f5ab0552f6a9

Observation d05ec6d2-877e-4df3-a98f-56f099406ef1 · outbound

This paper cites PyTorch : An imperative style, high-performance deep learning library.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PyTorch : An imperative style, high-performance deep learning library

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:02.180789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.297355Z digest=sha256:4a31e4a962c83eeaeaa65694d0b9ea51ef33d06eadbc59689d8b3eb40213a895

Observation d5100856-34e6-483f-8383-9210abf52b3d · outbound

This paper cites Your Transformer is Secretly Linear.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Your Transformer is Secretly Linear

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.300035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.300035Z digest=sha256:73fa1b457a0281d240d13b13458b085ea97c79d949a2c04594e084f331edb4e1

Observation f446cae9-71a6-4b8b-99eb-2b2ff23675c6 · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.303327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.303327Z digest=sha256:6f834bee92d9547d1b5f11cf60516ff78c3f5a8536727cea3b6115a3803d1170

Observation e5ecc08a-e6ac-4eec-84ef-bda8409605dc · outbound

This paper cites Societal Biases in Language Generation: Progress and Challenges.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Societal Biases in Language Generation: Progress and Challenges

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.306566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.306566Z digest=sha256:da7e7e4ccba85bb8797c8222b5641d319de4d525034011c0a6b2a36ece0c7434

Observation a73bffc1-c291-4e1b-b4cc-510722dcd859 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.309979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.309979Z digest=sha256:b9133464d375b7b208b10ee328b371d24c78dd1a4755e17c3822abfaa18ffd30

Observation a9b5555b-cc54-48a1-87a2-84fa96ffc1e3 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.313944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.313944Z digest=sha256:0850448a431ab27aab1c030e31ea0a26064da36179ea4c08a72ddc3449984cec

Observation fd55ea0e-1543-4e06-87e7-f2ab5767f297 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.316831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.316831Z digest=sha256:5ad2b4392218c761afbeaac7a6f7dea9f9406a37ccb55c5854c5b67bf3c897f4

Observation 9926d273-81d8-46b8-9d2d-605d88f75458 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.320318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.320318Z digest=sha256:883d83893c8327da52ca569a86fbbf8a6f7291e7f535ab48daa0daf4568fdfc2

Observation 5c3e592f-d489-49e3-825a-8c282473eae6 · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.162672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.323403Z digest=sha256:cb47cb7aed14ae84e1995006d0f409bf4a8de26c51defd97d9087b23668d372b

Observation f89d223d-34de-4a17-bc3c-b7f9018f125b · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.150601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.326506Z digest=sha256:93461cf55350bbf5d92d7eef69eeec99e4e00b7831826c7900190957dd70b6ce

Observation 9cefe2c7-1865-4dcd-b5e4-bdf3cb2700a8 · outbound

This paper cites GPTVQ: The Blessing of Dimensionality for LLM Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GPTVQ: The Blessing of Dimensionality for LLM Quantization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.330531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.330531Z digest=sha256:a71820fc392a5004cf2fc445a27d04f4d1d0d931fd8715f0506caa9b1fa39b50

Observation b1a1c0d2-a727-469b-95c0-61804011d507 · outbound

This paper cites Attention is all you need.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Attention is all you need

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.335869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.335869Z digest=sha256:2795897f6ecb5da91d286f914a41ac6e95f993c2c135f623961d6c26fd418f3f

Observation c20bd02e-0937-454f-a476-5e06d12ab2bd · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.339571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.339571Z digest=sha256:b0d0749d3e076ebed00eeadc1285733f05b3ffb3d3163998d2f02d894ddbb9c1

Observation 35ce49bf-5326-4b1f-8bab-7cc0fa564058 · outbound

This paper cites Ethical and social risks of harm from Language Models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Ethical and social risks of harm from Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.342390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.342390Z digest=sha256:7cde701046394c531e9fff7ae5461d2153386fd7592a49970b60a55fd2ddf9d8

Observation 27820dda-1aaa-41e4-b023-1620dc7e8ae4 · outbound

This paper cites an unresolved cited work.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:02.128852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.346803Z digest=sha256:d2f9ca37a9702ef524d7834cc6383629060e25e2484e064eeb3834dd55d5c613

Observation f5ec474a-361b-46dd-98f6-707024ab02b4 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.350988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.350988Z digest=sha256:2e413d1aab91ccc5e7f291792f2a336cc227fb4df1efb7af558804dea86df2ef

Observation 2645100c-35ab-44b7-a89d-6b2d880f895f · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Efficient Streaming Language Models with Attention Sinks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.355022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.355022Z digest=sha256:305dc53220b8facc1b73b46b09663ac304342ae42fc9da5ead65b69c31cd2de8

Observation 685d429d-b916-4c2d-b21e-3e38f7f79ca4 · outbound

This paper cites Qwen2 Technical Report.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.358141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.358141Z digest=sha256:dfeee848d5e743718f9ffee40dce2083becfcec542f85a54bc11a3564460ebc1

Observation e36af4c4-b7d4-4940-8b8d-a614cfbfadfe · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.361362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.361362Z digest=sha256:500add6f06f25c4775ddcfcf197e074423a12050c74f9faccf04d8aaa6e43cc2

Observation e75cced3-0e74-4d9b-a114-f262531871d2 · outbound

This paper cites KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.364320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.364320Z digest=sha256:b0bf0eaed43317dbf063d3f027af48abb6b6d55db049dd419cd5f9d67f0a7888

Observation 7761c196-453c-4558-84f1-8e046a244107 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.368235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.368235Z digest=sha256:d9f7af4f4083e895a965be2a1a22318c585b84f35f74e759c2d7df90b33d4a75

Observation 09da41e6-ac57-436c-8e5e-4482df3496cc · outbound

This paper cites QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.371452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.371452Z digest=sha256:40dfa96087254a68982977f0a51a5f01ddf82a331cdec38e2fc808643c5e493c

Observation e6291d10-d012-4508-af3a-b78590459ccf · outbound

This paper cites FDC: Fast KV Dimensionality Compression for Efficient LLM Inference.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models FDC: Fast KV Dimensionality Compression for Efficient LLM Inference

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-09T20:34:01.426422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.375678Z digest=sha256:fbedd4834452c9b519b320ff0a4f83b93e3656a17c9c78f8c9aa535f390d9333

Observation b458603b-e590-4201-bbfc-0d58d785b0d5 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.379174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.379174Z digest=sha256:22659caa383436f1beffd53f59ac2d8605dac62ae6a54042f5c13dc352b9c31c

Observation b11442c7-e928-4a77-9eee-1621558020bd · outbound

This paper cites Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.382301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.382301Z digest=sha256:09a6815e50ebcec248cf24818cd36f7ce6a8e281a89513808e1c56567fec8718

Pith citing papers

Observation 0c036e79-b643-41f2-aef1-56bf2295a7f7 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.037853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:49:13.353700Z digest=sha256:aa5c5ac6b8a27e267deefe2226268efae38a5806dc9feecaf1b113eda9380b39

Observation 1c7f9c93-c852-478e-a569-06591cefc921 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:56.308454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:48:21.341462Z digest=sha256:5a8091c57589671440d0fbe1afbd293dab721adf0fffc6ec0a855cb37bb95046

Observation 7779274f-0729-4706-a7d2-54b82190afe9 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.392954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:35:30.155592Z digest=sha256:99b4d95480fa808b1c87c69324617235126a5d8c3812415f30734fb7ab4d6cf7

Observation 34b564f1-c187-439d-8388-da7bf1f7fea5 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:42:39.675232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:41:30.021302Z digest=sha256:762727a207981c065216fb84bacead39a2706ed57b416e7a264e2255e36351f3

Observation 37f7136f-56d8-40a3-b40a-157028fe8598 · inbound

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache cites this paper.

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:33.526717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:16:15.869198Z digest=sha256:77b1ba0bf0c51b6535df624052b9b809a37b30a3699387ec820ea69db17ad29e

Observation 074dd5fb-7213-4fb4-8bac-8cce632be527 · inbound

A Simple Plug-in for Improving Eviction-Based KV Cache Compression cites this paper.

A Simple Plug-in for Improving Eviction-Based KV Cache Compression Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:24.206757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:51:04.068354Z digest=sha256:b11f76a80674051ee9e4a575e0bba1e2a338391d6d6e0ea6aa68a651d09123d0