Pith. sign in

Paper Citation Record · LEDGER

You Only Cache Once: Decoder-Decoder Architectures for Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2405.05254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05254 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:13:38.329919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.070431Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 76089141-4f48-4e51-8561-9c31cf68f0cb · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.415000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:c8b0b0f3e406d8658dc93c54e55abfa192afe8d52eff778f5bb9a57131752f33

Observation 0082e2e0-a6fc-43d5-b057-138f82f37d4d · inbound

Gated Delta Networks: Improving Mamba2 with Delta Rule cites this paper.

Gated Delta Networks: Improving Mamba2 with Delta Rule You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:50:24.320504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T14:50:23.991809Z digest=sha256:7aea79c44d36c87bc9372ea9a7b49974c93d091f7a059305ef75e30dca5aac35

Observation e9629a70-76bc-479b-b9ba-0bfaaaeab813 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.196543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:ed883a150ad017ca8c1f2f2da857cad18f80cb2b3959ef7253596c1ecae7f000

Observation 7eb8ef10-7404-4e97-b97d-a9d1fd8d5eeb · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.329919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.329919Z digest=sha256:40578d6fbd51d1b3b5de43d78b9921c8eb3709af7200a4d0c7649f5a2e143e1c

Observation ab040785-c0d6-42d3-94d1-0e5dd2b9d466 · inbound

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training cites this paper.

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:22.006655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:22.006655Z digest=sha256:521a39cf762912897f1a0728f05bc54c1e355ddeadd85d32e01a4fd1a991ed96

Observation ca85e05f-e502-4809-b022-bace30f01456 · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:27.714146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:27.714146Z digest=sha256:5b4ad565b140da69a5b6031f7dd0d31e4e2ba581795d0c00afa7a103c4cf5797

Observation a5db910f-b2cd-4b5f-be99-8f1a334ccd22 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.363103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.363103Z digest=sha256:c7e50600dc606d79cda8378ff57e391a5cdaf7cd86be6cb6e6af69b1772981b1

Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · inbound

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs cites this paper.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.234433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.234433Z digest=sha256:55c12546ba53c1d5c004edefe0aada2bf590646a58464d72c83de8ff1e539571

Observation 13304490-bca6-4a82-90c0-81c7b3e72b6d · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.771372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.771372Z digest=sha256:dbdbcfaa2bcaa786b7d786b6bd053cba508971ef964d2720326d62d33894bda1

Observation 6e83b91d-a9ee-496d-9d4f-db0645586b0f · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:10.824263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:4a6682fab8366e8090cbaf6de94ed230bc65d87abafebafab2c0f3e83ee00d79

Observation a2aa1d70-5671-4611-9f6c-ec2aff9c48e9 · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.458682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:05:25.465669Z digest=sha256:b6cfa30aa5c751534cc6ca889b067c4491eb92f7dd97cc356deaeee6296453ac

Observation 3cecfd45-a9ed-41cc-a1b3-4df3fd28b6de · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.083968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:10:23.224012Z digest=sha256:5c16d7bd9679c9acd1db7634fb842d8a75f1df8fb3f1168f53e0ad562a1722d2

Observation 14cd498d-80d4-44dd-8f87-f02c818b831f · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.419010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:6c4fb58a92aef70c8cb3358c8da5acc0c6fbdefb092764f4392ce304b49ad1fc

Observation 4f3f47b6-d345-4f36-9191-c7a620b9f0d3 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:23.547953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:23.547953Z digest=sha256:d48678fcf0356b2a512e9751fbdb4ceef1a3391bb13bdedcede6c3ffeb52864e

Observation 0a9fdbd2-e375-4bad-bf49-76840635598a · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.824157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:728d355a54f889970e4b7ae399d0ada6f5599eaf251d95c17788f5fd08243536

Observation 57cf74bd-3e17-483f-acee-f328c61f57d7 · inbound

Q-Delta: Beyond Key-Value Associative State Evolution cites this paper.

Q-Delta: Beyond Key-Value Associative State Evolution You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:26.548538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T18:30:51.523567Z digest=sha256:a1c57ad09a2545394922b492aef6be0f3509adddb576b6294788349a3cd14718

Observation 93396da0-cd30-4d86-b9da-056eb7a98255 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.071688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:e1b1fa59cbc031c3819dfd457354f23286a2058fc2030e17ef67e6d0b0a9405d

Observation 2348b8b4-20be-4ce9-ac7f-1ae1c7839e96 · inbound

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers cites this paper.

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T23:22:31.614891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:22:31.614891Z digest=sha256:9d2d67697dafd93cf54448defa071d89ff5b6824a6855aa4c919d6b0fac51f95

Observation 72683e7f-bb28-4eb5-b906-475de5930321 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.340861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.340861Z digest=sha256:6e7f3a18e71d0f0f424fc6a617869f4bfd021c4d72d3be73ff8d809353a75565

Observation eb06958c-9289-4e84-b333-3ac60a66a525 · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:45.924478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:45.924478Z digest=sha256:fffc112392314c27b4700aa0e63b8f02453909fbc5d135fe4f2fc732e78cc276