Pith. sign in

Paper Citation Record · LEDGER

You Only Cache Once: Decoder-Decoder Architectures for Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2405.05254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05254 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:13:38.329919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.070431Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 76089141-4f48-4e51-8561-9c31cf68f0cb · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.415000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:a5f7c8f7bb0116806b93b02319e31a4c5e5b99d815e40012b792ab17f3ca92a7

Observation 0082e2e0-a6fc-43d5-b057-138f82f37d4d · inbound

Gated Delta Networks: Improving Mamba2 with Delta Rule cites this paper.

Gated Delta Networks: Improving Mamba2 with Delta Rule You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:50:24.320504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T14:50:23.991809Z digest=sha256:f86345bb2e6b4d17eec93111f5089e924b94e57587cd526f24cedb2d8d557bf2

Observation e9629a70-76bc-479b-b9ba-0bfaaaeab813 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.196543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:d2011de78825c05d4eabdfa7d9f4f66054a6cbebd7e8574c11071f4fd663dea3

Observation 7eb8ef10-7404-4e97-b97d-a9d1fd8d5eeb · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.329919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.329919Z digest=sha256:40578d6fbd51d1b3b5de43d78b9921c8eb3709af7200a4d0c7649f5a2e143e1c

Observation ab040785-c0d6-42d3-94d1-0e5dd2b9d466 · inbound

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training cites this paper.

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:22.006655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:22.006655Z digest=sha256:3fdd8947ad7adcbbf92d793e45e67ce8a987919fb8c9894a76d4e980bd6587e7

Observation ca85e05f-e502-4809-b022-bace30f01456 · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:27.714146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:27.714146Z digest=sha256:5b4ad565b140da69a5b6031f7dd0d31e4e2ba581795d0c00afa7a103c4cf5797

Observation a5db910f-b2cd-4b5f-be99-8f1a334ccd22 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.363103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.363103Z digest=sha256:14e974b9266b68fd7408b672372608c7e694914d6fbef673e9d0218004dd8f8d

Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · inbound

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs cites this paper.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.234433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.234433Z digest=sha256:f9846fa251fd3d8f0505a67cb9b4cc207da57c203bdce86b980255816394e7fd

Observation 13304490-bca6-4a82-90c0-81c7b3e72b6d · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.771372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.771372Z digest=sha256:dbdbcfaa2bcaa786b7d786b6bd053cba508971ef964d2720326d62d33894bda1

Observation 6e83b91d-a9ee-496d-9d4f-db0645586b0f · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:10.824263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:2f5fec844dd655a1eb9706cad1d7096c1ca83c8ca019b5e54aa17f31d3364abe

Observation a2aa1d70-5671-4611-9f6c-ec2aff9c48e9 · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.458682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:05:25.465669Z digest=sha256:99b75837b75adb68f74a201e1125d063a84733e11b7749288a43609383cd19d7

Observation 3cecfd45-a9ed-41cc-a1b3-4df3fd28b6de · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.083968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:10:23.224012Z digest=sha256:67b8cb3177f9432af8ca9e1547b6bb74cc6fd863baaa6d08ba5fe7a793d7d777

Observation 14cd498d-80d4-44dd-8f87-f02c818b831f · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.419010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:8f0f0660199b799888e57b4a0824c8f40bfc12781489b9bf3549567bdb8de897

Observation 4f3f47b6-d345-4f36-9191-c7a620b9f0d3 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:23.547953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:23.547953Z digest=sha256:d48678fcf0356b2a512e9751fbdb4ceef1a3391bb13bdedcede6c3ffeb52864e

Observation 0a9fdbd2-e375-4bad-bf49-76840635598a · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.824157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:b387b20d5eff0f62c182007aa39684e72feba6c26a4dfc2080bddf34a89b1ece

Observation 57cf74bd-3e17-483f-acee-f328c61f57d7 · inbound

Q-Delta: Beyond Key-Value Associative State Evolution cites this paper.

Q-Delta: Beyond Key-Value Associative State Evolution You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:26.548538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T18:30:51.523567Z digest=sha256:7ef3de7f7b86596347a3152ff596e25ead52c130c4d6eac9ae4d73aa9b8d28ab

Observation 93396da0-cd30-4d86-b9da-056eb7a98255 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.071688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:c195fd9ab4279d7923b4669e22df336be66591481a613775718b3345422d0fe6

Observation 2348b8b4-20be-4ce9-ac7f-1ae1c7839e96 · inbound

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers cites this paper.

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T23:22:31.614891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:22:31.614891Z digest=sha256:a8327fe3e7d186d0d70a402e4b1e249f2ab3a2905cb50012b1682018c3eea1d7

Observation 72683e7f-bb28-4eb5-b906-475de5930321 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.340861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.340861Z digest=sha256:6e7f3a18e71d0f0f424fc6a617869f4bfd021c4d72d3be73ff8d809353a75565

Observation eb06958c-9289-4e84-b333-3ac60a66a525 · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:45.924478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:45.924478Z digest=sha256:fffc112392314c27b4700aa0e63b8f02453909fbc5d135fe4f2fc732e78cc276