Pith. sign in

Paper Citation Record · LEDGER

MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2406.09297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09297 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:13:55.342435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T20:15:58.175553Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b3ed83cb-0398-441a-9e59-5dc460e7eb44 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.342435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.342435Z digest=sha256:1db07a8ded9dbf044377da6736650eff497e91195fb665995801749d30fe8d15

Observation 379f3c06-11b5-4a3f-8aa5-1a0921f3d02b · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.840582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.840582Z digest=sha256:c20d0316dfc6ebcc6b5d99b15fb065eea34d670c63d4936fcd4fbe5f91b6b190

Observation f6a75ec7-d879-4c79-ab6d-85c446259519 · inbound

Multi-matrix Factorization Attention cites this paper.

Multi-matrix Factorization Attention MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.049741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.049741Z digest=sha256:1552562dfca172218ea07c563718bd452fbcbb3d7b5e6489617922c2abcba9a2

Observation e4a93f0a-d1ba-443e-afbe-23691d6d1013 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.826304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.826304Z digest=sha256:0fc5959ffc49d0fa86b29a9ba25bf199f7beb9a97dbcaed70b69c5c7fdd9c1e9

Observation 8c54e9f0-2323-4fed-adcd-d76970a55959 · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.208529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.208529Z digest=sha256:9cf0c21f601e9baa8513e0db7b6706ee7bc57ca69dfd07257b8e7bcfd1b57e62

Observation baf7dfb4-9da4-476f-9c9b-318038147008 · inbound

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures cites this paper.

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:15:58.298505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:15:57.782754Z digest=sha256:3c1850b698f6017bd80a9a1f8628619fa9337dd7f4a3775d5529ece52225ab89

Observation b9024bee-7640-4399-b436-2a17ca8b25cd · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.347596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.347596Z digest=sha256:6354b5c1e854ec3aa5be2fa09a3a64ba082232c706f8d5555d162d98c3672f9a