Pith. sign in

Paper Citation Record · LEDGER

MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2406.09297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09297 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:13:55.342435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T20:15:58.175553Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b3ed83cb-0398-441a-9e59-5dc460e7eb44 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.342435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.342435Z digest=sha256:de2b523bac224ee78c65d4f4d17eac58ed89d2894fd61e61a1d7686693ca351d

Observation 379f3c06-11b5-4a3f-8aa5-1a0921f3d02b · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.840582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.840582Z digest=sha256:98540c620e4fd5c899ebdb2323fc447312105ecc073f98dc2728bda1a8021fac

Observation f6a75ec7-d879-4c79-ab6d-85c446259519 · inbound

Multi-matrix Factorization Attention cites this paper.

Multi-matrix Factorization Attention MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.049741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.049741Z digest=sha256:0b6ce9e59ee361af037aa6192bd40a9c77d662ed1a0c3055262f5820fa3164b9

Observation e4a93f0a-d1ba-443e-afbe-23691d6d1013 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.826304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.826304Z digest=sha256:33f379629052df902be8fdff4338231998c98aba353f6f04d22b7692e5b3bb3c

Observation 8c54e9f0-2323-4fed-adcd-d76970a55959 · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.208529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.208529Z digest=sha256:93521396bb0b0b552c0f59ff231dc2bd1d25faa95074c043f7ec7f7e66247c0f

Observation baf7dfb4-9da4-476f-9c9b-318038147008 · inbound

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures cites this paper.

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:15:58.298505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:15:57.782754Z digest=sha256:0ac305f9a01f409047852b2afd88432324f0707b161b8e5978cc74bddb8432a9

Observation b9024bee-7640-4399-b436-2a17ca8b25cd · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.347596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.347596Z digest=sha256:80c9b3c080c151e7d953c2474bd2803386bb292f4268670963bbc2808d80e752