Pith. sign in

Paper Citation Record · LEDGER

Palu: Compressing KV-Cache with Low-Rank Projection

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2407.21118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.21118 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:40:37.603820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.121083Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f73c71fa-3690-4967-8845-19cad3674fec · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:49:33.802941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:0a8d24353275218a46be711950d9caaf293725a34152be68dc744951647859ce

Observation e3967fdd-1b7c-46d2-ac7e-5fb0b7207041 · inbound

Attamba: Attending To Multi-Token States cites this paper.

Attamba: Attending To Multi-Token States Palu: Compressing KV-Cache with Low-Rank Projection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:33.034655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:56:33.034655Z digest=sha256:e7f5b79fd21138c64839f76aa2ca4f22201f5e16a39924a66725b84adab0bd3b

Observation 323c0ab9-4f1b-4b74-ad4e-9bb33b0d543a · inbound

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals cites this paper.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.776013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.776013Z digest=sha256:69d39acb5fe85fe1c7255d41290e3e60add86027a2b5070ef4bcdcfe42aa4908

Observation 3ff7df90-6c6e-49d1-80ec-2aa6f8355212 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Palu: Compressing KV-Cache with Low-Rank Projection

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.985213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.985213Z digest=sha256:76baf55ad858bae008eb1ba4aaaced7518a258f269d686a87e32b1546dcb2590

Observation 9e72fe70-0f02-408e-ad53-19ad07a7a089 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.189891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.189891Z digest=sha256:fdae7edb3dfd0c927e6b8117c905e846fa4a33e33bf06b1d810b71b0a1837d36

Observation def58ec3-b6a0-4d23-a0ba-3f4148078342 · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need Palu: Compressing KV-Cache with Low-Rank Projection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.171526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.171526Z digest=sha256:4193d0edffa5ade1ebf5280658c6354c2302b17ce78792db7ba72ef39460e04d

Observation f857c088-56f1-40a0-b779-0614274d8b18 · inbound

A3 : an Analytical Low-Rank Approximation Framework for Attention cites this paper.

A3 : an Analytical Low-Rank Approximation Framework for Attention Palu: Compressing KV-Cache with Low-Rank Projection

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:04:57.131560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T15:02:33.864861Z digest=sha256:5266b973fca0c416afde46dad21e04448d839f0af6ab4f34bf09fc4180160d4c

Observation de377a09-f03b-457f-b1bd-0eb0dffa5fe3 · inbound

LatentLLM: Attention-Aware Joint Tensor Compression cites this paper.

LatentLLM: Attention-Aware Joint Tensor Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.552150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.552150Z digest=sha256:c5721411dc362b00971507f8087b779a95417ad3b6f53d35d65fd3f894de2733

Observation 7e5ea5fb-a49e-40fb-b73a-2d4c9cb73d6a · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Palu: Compressing KV-Cache with Low-Rank Projection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.031025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.031025Z digest=sha256:c438770cbae87c129652906fe50c17491adc4be1cb26e183584c7c32f1ab7c3a

Observation 484ef146-5291-4782-9dd0-fdc49ba24a4b · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.825323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.825323Z digest=sha256:abe92dcc41ca2793bd1f13f6b3f570961b620b311284146d4db78f85c02ae1b6

Observation 6b808613-58d1-4cf3-b466-da47f926ab65 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Palu: Compressing KV-Cache with Low-Rank Projection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:26.941477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:26.941477Z digest=sha256:6d72d247719c8fbf2442e4d84bd49a1ee10455110d66e19282b15f2bcacd3ef7

Observation 77689f63-3157-4247-bc97-f1a0c49a8c57 · inbound

Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache cites this paper.

Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache Palu: Compressing KV-Cache with Low-Rank Projection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T01:10:37.292088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:10:37.292088Z digest=sha256:df23c976b5d651e36f6e50b76e045023dcb1a58fbc58410b65537f0f2bd98366

Observation 11c2edea-ccd7-4683-aea8-76ea40ffa21c · inbound

RCStat: A Statistical Framework for using Relative Contextualization in Transformers cites this paper.

RCStat: A Statistical Framework for using Relative Contextualization in Transformers Palu: Compressing KV-Cache with Low-Rank Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:40:37.603820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:40:37.603820Z digest=sha256:96b00680178fb4b01d5d462bff9bc319ffd93a6df35b45d283693ce8db09636b

Observation 72c9b27c-2c3b-4a5d-87d3-bc691b0d4124 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.670799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.670799Z digest=sha256:ea73002f2235119565fa155ab0286bcefc880df6d06fed308dc343a5bacc6a5d

Observation f929d593-9219-4713-9a8b-eef18834e01c · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:06.348432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:06.348432Z digest=sha256:0e408f52e5eee6bc60f13736514b9c77b983ec166d4a250ee7f4ed7a598b6291

Observation 98b382e6-6302-4125-9fce-9adea19682cf · inbound

OjaKV: Context-Aware Online Low-Rank KV Cache Compression cites this paper.

OjaKV: Context-Aware Online Low-Rank KV Cache Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.651969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T13:26:02.980973Z digest=sha256:c6ab778697abf821ca30eb11c0fb4b55186303fdb8021566d1d3bbde68307096

Observation 57baefaa-2193-4053-b8d4-b8587bdf6c13 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.971195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:469ff84cd479efaa2e1d23ba6a21f6d81975f96fcb3610df252acfb47067196c

Observation 7ea17c51-bd2f-45bc-9051-11e8fed10732 · inbound

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models cites this paper.

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:18.020672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T21:05:09.254086Z digest=sha256:a48b7afefdc2346fc5297834b02de00162b79b28907ec7d1a2e3e112650f689c

Observation e152ac02-d2c4-459e-8575-17a5922a051c · inbound

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization cites this paper.

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization Palu: Compressing KV-Cache with Low-Rank Projection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:50.983412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T20:18:04.392331Z digest=sha256:02b09a9deb36b6fb9eb9b8153c465d0e36c0d5c2516ab8c13d52782ae8a492aa

Observation f5475433-8212-4511-bf65-d1c982f6ba30 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization Palu: Compressing KV-Cache with Low-Rank Projection

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:18:16.527054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:a4a1c76f5e1edc87a19820aea3d80bde22aa777f7d9ee3dc81ea6bd8f78f0c7b

Observation 9ab8b0ee-4d6c-4b4b-b1ec-b4cf79efc620 · inbound

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference cites this paper.

Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:43:23.770713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:43:18.828740Z digest=sha256:0c69ff8ee9b5ed5b3f81ccc509a7de833d743d39a05685b5b17fa6bccf1fd9f9

Observation 554a9a90-546c-4240-a260-d11ac706e328 · inbound

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs cites this paper.

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs Palu: Compressing KV-Cache with Low-Rank Projection

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.774202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T22:54:55.101568Z digest=sha256:a0d74fb52b149930c9c5b476077507365019bd76ae76f223cdff1c2e7d13e28b

Observation 06d2376f-b0a2-4a6c-a291-d84b8421dc2c · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation Palu: Compressing KV-Cache with Low-Rank Projection

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.508413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:3f2282702259194596d9acae937a29af9f8a24cfb0b93f9ca6934a06d0db1521

Observation 90c87189-d919-4472-a996-0f980e90c98a · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.707197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:c2b2755eaace6144c6e4beba7130921ae6f200204a71dc6958670ddb54f5a142

Observation af4941be-b6f9-4ab6-80a6-1de7e319cc81 · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM Palu: Compressing KV-Cache with Low-Rank Projection

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.699258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:f88e556db54e87c8bb7d780c3b9bdc4f0630b0ac7cb36669fa15e3db6d93f768

Observation 69977081-c482-4a04-864e-82f935153758 · inbound

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models cites this paper.

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models Palu: Compressing KV-Cache with Low-Rank Projection

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.348156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T13:09:37.979051Z digest=sha256:0b291a4f2f56ee4eef6c49fa6c9f37ef7d6376c54e7c74940e19fe77893b0522

Observation ec9e31b7-358a-4438-9a9a-195ca299eee0 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Palu: Compressing KV-Cache with Low-Rank Projection

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T01:36:44.122209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:928083d639124eddcac359c6c06480ddeea878d619d4f3d8f59b69a34cac1567

Observation 3e2573aa-faa8-4189-b696-6d4092c86a7b · inbound

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs cites this paper.

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs Palu: Compressing KV-Cache with Low-Rank Projection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:31:50.083575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:31:50.083575Z digest=sha256:617b2c6200fdee5f232dc7fdbf49ed9942d96c346df5768682d381f454b1073a

Observation 7c90bfd3-a5db-4c70-8f38-0575b9e87fb7 · inbound

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge cites this paper.

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge Palu: Compressing KV-Cache with Low-Rank Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:50:14.941014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:50:14.941014Z digest=sha256:57f4d7e3c642f1000369ce82b43a17922e19795593bda16a4b1930f6de92a5b5

Observation d4431f5f-d32c-4e5f-9a5b-6ed319e2395e · inbound

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation cites this paper.

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation Palu: Compressing KV-Cache with Low-Rank Projection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T18:02:37.410810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T18:02:37.410810Z digest=sha256:5512c4ab03c1f705531ad25c946388761d4dd24fa7b974f9c7ae8062c7ddf300

Observation 9efb6cf7-0d03-46c7-bd1f-c19c47410427 · inbound

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding cites this paper.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.614843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.614843Z digest=sha256:697538c7f003b647df491dbbe6e7be0388319bb908a870c1ebe9adf37c28b7c5

Observation 0cefdd22-aad1-4e4c-80c1-cb14334d7f0d · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:48.033830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:48.033830Z digest=sha256:eab106cb1c738ff2607cfbab24f782a40d66ee065a8b0324bfa37adc6d502fda

Observation cfef5d0d-dce5-4308-b32a-06239cc816f5 · inbound

Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes cites this paper.

Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes Palu: Compressing KV-Cache with Low-Rank Projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T07:49:39.405147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:49:39.405147Z digest=sha256:9ef173af6e2e6150d01fb3b981b8aab706d0cdd7b7986cf0a8742c9ac3780dd6

Observation 064a8211-11ca-45ae-b1ce-7bd7c3b7bc12 · inbound

AnchorKV: Anchor-Residual KV Cache Compression cites this paper.

AnchorKV: Anchor-Residual KV Cache Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:52.922607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:02:52.922607Z digest=sha256:1ae6010c65f809636e4b9934c1b4f46c13500d6efb2ed6966a403e734d2d8130

Observation d3be1814-e7c1-4d18-aa49-b2d1acf584b7 · inbound

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval cites this paper.

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:35.285100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:35.285100Z digest=sha256:7281cd0265fa287b315b12d89d7e56f58303d41de9da3269647d004b21654c07

Observation b4baf7a2-63fd-442c-ba6c-a225c61f16dc · inbound

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval cites this paper.

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:54:42.694725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:54:42.694725Z digest=sha256:b1da49f83691a1d8f2a6947e7d1e39aa965813437040c12f841526791fd81691

Observation ebec126e-da94-419a-a482-322eae95ddbb · inbound

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference cites this paper.

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference Palu: Compressing KV-Cache with Low-Rank Projection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:46:03.872365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:46:03.872365Z digest=sha256:218fcae916c32e0f6ece3ba6f71c7068309ac6e34f853d51bd3e708caedd0d0f