Pith. sign in

Paper Citation Record · LEDGER

Training Transformers for KV Cache Compressibility

As of 3 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2605.05971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05971 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:01:40.766843Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:35:22.750886Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact23
  • verified fuzzy35
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbf67cf-30da-48db-9871-2b97be1dc410 · outbound

This paper cites Can Foundation Models Help Us Achieve Perfect Secrecy?.

Training Transformers for KV Cache Compressibility Can Foundation Models Help Us Achieve Perfect Secrecy?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.987477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:31082f051b1b6fbe8867280aea077289cfb4a515cad3c6446677760eaa9cea67

Observation aef4b139-e3c5-4177-9b93-3923e41c6c8d · outbound

This paper cites Longbench v2: Towards deeper understanding and reason- ing on realistic long-context multitasks.

Training Transformers for KV Cache Compressibility Longbench v2: Towards deeper understanding and reason- ing on realistic long-context multitasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.294421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7fefeb22089eab1282ab78136c6efa724f4a9dc5d8e048671dd5a36bc0fe3a58

Observation b4d5c7da-b73d-4ec3-a51a-c821f65cb593 · outbound

This paper cites Longformer: The Long-Document Transformer.

Training Transformers for KV Cache Compressibility Longformer: The Long-Document Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.990094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6bdb36a98d6f98b637806757d044f6b20ac36226fc8350532a5d0514cc950b6b

Observation 3493bf0e-5511-4291-a28e-3e97eaa32b7a · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

Training Transformers for KV Cache Compressibility PIQA: Reasoning about physical commonsense in natural language

Reference 4

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.793734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:951f4c7083f3a7efd47d1905686e736d3ed56b1d1102488a58fcb33f425d97b0

Observation 7be11d93-61c2-4889-b04b-0795f160ba40 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-13T06:02:24.289944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5e46717c5f9e0413881fabc129ab733a706e75f3479b2f9853b56f2d43e779ce

Observation 9936c8c6-164c-45d0-93f7-f81526078cb2 · outbound

This paper cites PyramidKV: Dynamic kv cache compression based on pyramidal information funneling.

Training Transformers for KV Cache Compressibility PyramidKV: Dynamic kv cache compression based on pyramidal information funneling

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.287922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a28ef459cfb0bb00cdeba412a1699833a840ef7d551a64f1a346520859aec210

Observation 66d1b34d-c5f2-42f7-8b5e-bc7920e8e19c · outbound

This paper cites Doc-to-lora: Learning to instantly internalize contexts.

Training Transformers for KV Cache Compressibility Doc-to-lora: Learning to instantly internalize contexts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.959138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:728522a34f507636f80504395f6de745d64bbcd46050cb4b56d691dc6630726b

Observation a658111a-bcc5-49ea-9152-d1f6504481de · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Training Transformers for KV Cache Compressibility Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.981495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ec32860e93caf691c3ee19447fca790181882b914e303649ec37ce0bce11bea8

Observation 1eba9c7d-de47-40c9-8bcf-90e758f08cbd · outbound

This paper cites Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass.

Training Transformers for KV Cache Compressibility Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.953195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5b2cbe9f078115ec8553407e92fe93f3bc685e53676601842b9a5b0a2fa9ba74

Observation 3c77b851-8b43-40e8-9fbd-2e4dda0c619b · outbound

This paper cites Adapting language models to compress contexts.

Training Transformers for KV Cache Compressibility Adapting language models to compress contexts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.285790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3597baf4f8ca6c8f4888c0319615950f5311a54d8c9084e3478ecd4a9f1a8b29

Observation d9c843de-99c1-4fab-9c12-28c8a18157f2 · outbound

This paper cites Rethinking Attention with Performers.

Training Transformers for KV Cache Compressibility Rethinking Attention with Performers

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.943990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:defcb5e678d170f2389b38f749007ef4a846983be01dc6e1fc1e7e95f833d381

Observation 4db62a8c-2dd1-455a-a48d-c04e2aa53630 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Training Transformers for KV Cache Compressibility Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.967996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:8c95503b2ffa9106354376e6e1bc65ac8126b56211a6cbd6e8bb985f09111e01

Observation f93f265d-ef37-4145-ac87-3a3901e4034c · outbound

This paper cites InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers).

Training Transformers for KV Cache Compressibility InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)

Reference 13

Resolution
metadata mismatch
doi, observed 2026-05-13T06:02:21.797274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:33c6de8475a33dede1109faf9e37da2e3ffc3241fbb76c849bd29f107408e8e1

Observation f84f24d0-bfe3-49f4-9a63-c1def112e4f3 · outbound

This paper cites Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314.

Training Transformers for KV Cache Compressibility Approximation by superpositions of a sigmoidal function.Mathematics of control, signals and systems, 2(4):303–314

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.278883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:f405766df069743a64b06694c7e89a41aecfa8fa830d6638e708fa6d87ef9144

Observation dae1ed25-d9e0-412f-9bfd-c30c20856816 · outbound

This paper cites The centered convex body whose marginals have the heaviest tails.

Training Transformers for KV Cache Compressibility The centered convex body whose marginals have the heaviest tails

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.946969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:c7e91561d629b3ea3e5f66bb16756259c0b78bd020c3aee764c1be86b56695a1

Observation 9e953d73-6031-42fc-9a30-726e1da7f055 · outbound

This paper cites Cartridges: Lightweight and general-purpose long context representations via self-study.

Training Transformers for KV Cache Compressibility Cartridges: Lightweight and general-purpose long context representations via self-study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.283639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:931c29323114b4c0d40bb358b50571ed279052dd3ee2b829a2702cd992c0bfe0

Observation b5a6d708-345a-4c78-bccf-c14a19e9d834 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Training Transformers for KV Cache Compressibility Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.997183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:1fb3cb24bcaaf378ce1e482a070fa43fd37936850540954ddb0bf8eebc8b1a9c

Observation fa307a39-d96e-414a-895a-048eb46a3527 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Training Transformers for KV Cache Compressibility Efficiently Modeling Long Sequences with Structured State Spaces

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:21.984721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6b8690f20a92163859bc396c41142700708ff214d4f4c2468b2f9fde65256f3b

Observation af3ae607-84bb-4fd7-bb85-7ca3fc615b42 · outbound

This paper cites Lighte- val: A lightweight framework for llm evaluation.

Training Transformers for KV Cache Compressibility Lighte- val: A lightweight framework for llm evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.281046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:06f2bf6040020004e20ffa6bdbfd9bf2a8a49672fec48ced4cf74e734befc451

Observation 90d9039f-cf5f-484a-8221-1c221834db3b · outbound

This paper cites Delta-net: Real-time network verification using atoms.

Training Transformers for KV Cache Compressibility Delta-net: Real-time network verification using atoms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.296958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:4c8b25d73380d682bfcaf3cab1f046b9de4edcac14a049eb022307ea3eb8aed7

Observation 7eec02fd-e121-411b-a8aa-89dbac8dd5cb · outbound

This paper cites Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257.

Training Transformers for KV Cache Compressibility Approximation capabilities of multilayer feedforward networks.Neural networks, 4(2):251–257

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.274052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:12201dde1aef2f10acad8bafdebb643b152fc98afb22640f14057ad7f6d5bcf0

Observation 68a00469-4c6f-486a-9f23-033449813713 · outbound

This paper cites Dynamic Chunking for End-to-End Hierarchical Sequence Modeling.

Training Transformers for KV Cache Compressibility Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:02:21.965287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7cd2929e7ece203cb4fda41ffd60871105ec8b5fb02da16bce062fff72fce351

Observation 9a326a7c-54ae-4f48-bf4c-838d0e8a0510 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

Training Transformers for KV Cache Compressibility FinanceBench: A New Benchmark for Financial Question Answering

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:01:13.018607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:1322725b213adc72d277e608d2d8bfb1a8520e1152c851e0693a5590a80d5e12

Observation 452ddc86-650e-4627-9d2e-738e064716b8 · outbound

This paper cites LLMLingua: Com- pressing prompts for accelerated inference of large language models.

Training Transformers for KV Cache Compressibility LLMLingua: Com- pressing prompts for accelerated inference of large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.271575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:06ed1705fe9dc72f1ae34f8b3a55b3969d1459d6400211708a8d498697dede40

Observation 147a14f9-4813-44e6-8fc8-36fa5f80ccbd · outbound

This paper cites LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression.

Training Transformers for KV Cache Compressibility LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.276579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:45fc540472a9585d3d89d95f4c7eaff8401093cdcf0084634dfab7d444e01c9d

Observation 7b6c23f3-f5ff-4d25-a284-f34b0ee6cc09 · outbound

This paper cites Optimal experimental designs.The Annals of Mathemat- ical Statistics, 37(4):783–815.

Training Transformers for KV Cache Compressibility Optimal experimental designs.The Annals of Mathemat- ical Statistics, 37(4):783–815

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.269453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:22c5a290c6f933d5fc002fe511c1886c95692d2abe1464cda661ebd432f01e80

Observation 7677d385-0243-47cc-85d4-c0dc0da6fdd1 · outbound

This paper cites Tchebycheff systems: With applications in analysis and statistics.(No Title).

Training Transformers for KV Cache Compressibility Tchebycheff systems: With applications in analysis and statistics.(No Title)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.264839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:0733147d33b0c2e23c90b5af632839929cdaec125ace4ca2d7c79ad7d6244e45

Observation 14256ac9-4194-45c5-81a1-37e532fa0702 · outbound

This paper cites Chebyshevian spline functions.Siam Journal on Numerical Analysis, 3(3):514–543.

Training Transformers for KV Cache Compressibility Chebyshevian spline functions.Siam Journal on Numerical Analysis, 3(3):514–543

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.267004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:60136b2361179c14dd76349c4a298a4eb73df4f378ac90cb1a2e24d357d08811

Observation dc366d46-17fc-4c89-9a35-3a056296534a · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Training Transformers for KV Cache Compressibility Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.260987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5ab6a4b0033beb5e3b4a52b88f9ac952cce139f5289ca2609cb52fec893e6fb1

Observation 265f9aee-8bae-4477-8744-baafd32ad3aa · outbound

This paper cites Kvzip: Query-agnostic kv cache compression with context reconstruction.

Training Transformers for KV Cache Compressibility Kvzip: Query-agnostic kv cache compression with context reconstruction

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.971227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:550d5c341f46ceb118d0a373046463ce7df4760d26deebb4cc6e36965289b143

Observation 0c3ab234-0bc3-47c1-a830-35e95aba6680 · outbound

This paper cites Lexico: Extreme KV cache compression via sparse coding over universal dictionaries.

Training Transformers for KV Cache Compressibility Lexico: Extreme KV cache compression via sparse coding over universal dictionaries

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.254451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:fc7c33b23fd29624fdbe3644b5225440ad4c7f7a7719e2fa720d8e3bfb9d8fdd

Observation d9c0c9a0-9b89-4abd-a86b-dc07dc7fb28c · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

Training Transformers for KV Cache Compressibility Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.256522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:fa556b94286982e6576010823132dbf479a19a4b0b823d8f834b87a03ce0036b

Observation b36181d0-9a4f-491d-9dbb-4544c41064c2 · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

Training Transformers for KV Cache Compressibility Compressing context to enhance inference efficiency of large language models

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T06:02:24.258780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6d90d1ecc480201084d16f3595369a06342d78b9293d15855e43ff0cdc13642b

Observation 0d344f97-1b95-470f-8156-b066ee485948 · outbound

This paper cites SnapKV: Llm knows what you are looking for before generation.

Training Transformers for KV Cache Compressibility SnapKV: Llm knows what you are looking for before generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.262982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:5abe318d582ebc59d671d1a3812b5043bb37391bea971adb3bcef3c4ae1d5f34

Observation ced269bf-c1b2-4443-959a-8fa0189b7ab3 · outbound

This paper cites LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence.

Training Transformers for KV Cache Compressibility LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.993912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:2e37f2c2636fd87ed3835916c148eb2645d018cfae6374078258f0345e177c9b

Observation 89401e05-302f-4574-9d6f-221366601686 · outbound

This paper cites SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass.

Training Transformers for KV Cache Compressibility SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:03:41.308564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:f50fddc628efe0b9bde526036d5bbf8d0c67aebb717f6d4af13ae97b42343118

Observation 628d164d-0462-476f-9b91-c95e952b6feb · outbound

This paper cites Pointer sentinel mixture models.

Training Transformers for KV Cache Compressibility Pointer sentinel mixture models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.292079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:6fc557e6330c636f838e91e7a56e88b74836925ee02b718b026b3e8aedda7846

Observation 8976502a-1931-45cc-a261-2796fd872c02 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-13T06:02:24.250037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:16133a734faacda8ac7bc9a4a5d5dcac809eea1713de72b07d60ef7abeb1dd74

Observation bd3951d8-cdd7-41c6-9b01-43f233d1d814 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Training Transformers for KV Cache Compressibility Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 39

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.790714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:1cfcfe6ade957698549dc77f28ab5c8bd9adc5b215951d785799d72f14bf05b3

Observation e313e9fa-7d48-4e25-b6fe-6b3a50bef301 · outbound

This paper cites Learning to compress prompts with gist tokens.

Training Transformers for KV Cache Compressibility Learning to compress prompts with gist tokens

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.235100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:d07a172481128b3d60e2842055c6a76a3551087688e506e916355575b3dd0fa9

Observation e5718b8d-1ccb-478e-8d9b-d9a69704d022 · outbound

This paper cites Using an llm to help with code understanding.

Training Transformers for KV Cache Compressibility Using an llm to help with code understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.220818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a18abe2d1a830b8485a2b1cbcca3c0ddfadb19d8e5534a0a0fdf007e6c1c0f27

Observation c076a8b2-29b6-4f6a-a674-b13ac4d9bd53 · outbound

This paper cites Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025.

Training Transformers for KV Cache Compressibility Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.223236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:0a27baa981bd6acef86a4604368a3a121b6665b44af068c7bef365fb977f615c

Observation 436b9803-9b50-47e7-8d5a-f49bbe585cbc · outbound

This paper cites Transformers are multi-state RNNs.

Training Transformers for KV Cache Compressibility Transformers are multi-state RNNs

Reference 43

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.803106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3a3a29c15793561de0689b93e6569d9e75e27315fd75920e8c4f4a83325c78c0

Observation 393bdcbc-4be4-4d49-99d4-95bf24436637 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849.

Training Transformers for KV Cache Compressibility The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.227937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:0d59f6fc1c99e60de97a10e7f78d6270776cf383a9d87b58df9a1acc3e94cc5c

Observation 4c76dc23-6216-4dfa-b833-1a6b6a2f153f · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Training Transformers for KV Cache Compressibility Compressive Transformers for Long-Range Sequence Modelling

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:46:17.004740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:58e0ecf445c76830709272ce5177311b1df8b198c7ce7a3ea3f61417bbba8cce

Observation 40abb738-a348-46b4-aa94-045157a1adbb · outbound

This paper cites Effective context engi- neering for ai agents, September 2025.

Training Transformers for KV Cache Compressibility Effective context engi- neering for ai agents, September 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.218665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:8dcba5a0fdeb9fda7ce6c55322ad775555d4bbface74f1927aeccb7d0b910410

Observation 59cfac00-a8e9-4f9b-a166-4993e0673cf3 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , author=.

Training Transformers for KV Cache Compressibility Proceedings of the AAAI Conference on Artificial Intelligence , author=

Reference 47

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.786005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:d208f57f51b689a869e3bec1744e69bc27a3f06dd2f6b4b36ae7cfd759f432f6

Observation 1d3aa75c-7184-4ffb-8779-6622470094bd · outbound

This paper cites Representational strengths and limitations of transformers.Advances in Neural Information Processing Systems, 36:36677–36707.

Training Transformers for KV Cache Compressibility Representational strengths and limitations of transformers.Advances in Neural Information Processing Systems, 36:36677–36707

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.213905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:bc3864280e1236295caafd46b414cedf5be4fd55f0f089cc3023c22c0d33b832

Observation fe83771e-4d4a-471c-89e7-526400e0d2da · outbound

This paper cites Social IQa: Commonsense reasoning about social interactions.

Training Transformers for KV Cache Compressibility Social IQa: Commonsense reasoning about social interactions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.225530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:315cd3f85d707c45ae443bd0f0921e2f3a19919890f55f79776a6bcae9ab1890

Observation 3f357129-8531-4ade-8248-1d2c178f607b · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Training Transformers for KV Cache Compressibility Social IQ a: Commonsense reasoning about social interactions

Reference 50

Resolution
metadata mismatch
doi, observed 2026-05-13T06:02:21.800181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:428b43551ce03a3ef9554c3ff577e22be1f8c65c2ea7c7c8af7d9f9f60e3af1d

Observation ce7ad80b-0f42-46bf-8ae6-71a8d92048bd · outbound

This paper cites QUEST: Query-aware sparsity for efficient long-context LLM inference.

Training Transformers for KV Cache Compressibility QUEST: Query-aware sparsity for efficient long-context LLM inference

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.252230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:955f3eb3659c4499d14aeb21adc769b35040c88d4973218a06908006be6eca38

Observation 3080a2c0-cb92-4ffd-81ed-a3fcae3bda5f · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Training Transformers for KV Cache Compressibility Qwen2.5: A party of foundation models, September 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.204011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:a2e9702ebdfdd892f405237de6f045b6aa77ed61168be5422429348b05bcade6

Observation 8a2a4e3b-5597-493d-bd83-ac4725c3f36a · outbound

This paper cites Efficient streaming language models with attention sinks.

Training Transformers for KV Cache Compressibility Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.208978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:cd1f8560cdf1d1db1a19472f011373a190ad32606747e51a110e4a7eaf131137

Observation 0bd4eb5c-44b2-47f3-b70f-859d0874db08 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Training Transformers for KV Cache Compressibility DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:49:16.947404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:da1b135e6829de6e7ac9ec5f0a56e989ac43b50939a9be6ac98f0b3b29895067

Observation 6113bfcc-ac39-448c-a9d4-21b8a1d1bba4 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Training Transformers for KV Cache Compressibility Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:50:25.025022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:c931a3fd08a51bc95e3cb40655104116a444f83073a09faa5dd0eb748ec8ef0a

Observation dc536cf1-3d8d-4043-b146-f4033d7e12e7 · outbound

This paper cites arXiv preprint arXiv:2407.15160 , year =.

Training Transformers for KV Cache Compressibility arXiv preprint arXiv:2407.15160 , year =

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:22.000080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:7cdb672277c1ff4b2d34d0d48ab0687a61f72c047c1ce73eaef0591c028ac3a7

Observation 2bb92f07-362e-4e28-b0da-f1f7bf28eabe · outbound

This paper cites Deep sets.Advances in neural information processing systems, 30.

Training Transformers for KV Cache Compressibility Deep sets.Advances in neural information processing systems, 30

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.237288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:87b38e1d163fbc09edcd3fc1e973297501aacdb5a1828f089b5aaea01db1297c

Observation 5cffa41d-2c43-4d58-931e-df3ed14a0102 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297.

Training Transformers for KV Cache Compressibility Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.230077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:97ccd0a0b403d10dc55e258572128f13aa0f8c36e08cd1d905df1ac409f7a99b

Observation 6557247c-93a1-4764-9f98-24dfbadeb32a · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

Training Transformers for KV Cache Compressibility URL https:// doi.org/10.18653/v1/p19-1472

Reference 59

Resolution
verified exact
doi, observed 2026-05-13T06:02:21.806038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:023de273c44905c21febfd3d774b53c930274a8f7271117b434abab6b458f410

Observation 78c43fa5-9a10-49a3-84fc-d3d4cf794a20 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

Training Transformers for KV Cache Compressibility H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.216232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:ef4b3f7b706d98563e495203410b094beec87d07b6fc662771f3b06b5fd4e9fd

Observation 7444f49f-cb4e-4adb-83e3-5f870f1a0af6 · outbound

This paper cites Lifelong learning of large language model based agents: A roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Training Transformers for KV Cache Compressibility Lifelong learning of large language model based agents: A roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.232644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:cb624e060efe568737ead0c34224959cc0e1ac205e68c57a1ce4d5edebafd5ce

Observation 7b4e5337-878f-4ec0-a002-6601849bdb1c · outbound

This paper cites Fast kv compaction via attention matching.

Training Transformers for KV Cache Compressibility Fast kv compaction via attention matching

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.244984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:b7a789caf03def5dc6f3978f50e9f64b6ed5bdb05841c447cf8dfda1db685d53

Observation c6758ec7-0854-4dcb-8ee1-936b6165bf6a · outbound

This paper cites an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(70).

Training Transformers for KV Cache Compressibility an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(70)

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.247803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:1f972df0469e0f434c66c69b4db2391ba346d2503f14efc604349b56b9e7553e

Observation 65356efd-af18-492c-b640-be94d32de4c7 · outbound

This paper cites an unresolved cited work.

Training Transformers for KV Cache Compressibility Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-13T06:02:24.206400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:e64f210e30af286d918ecf595883a10e4ebba80922980379897073ab0de71e81

Observation 7ac604d7-f5e5-41b1-a5f0-4b3f2221255f · outbound

This paper cites an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(97).

Training Transformers for KV Cache Compressibility an)∈A n withn≤N, ∥f(a)−M(a)∥< ε.(97)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T06:02:24.211502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:295747945c9f38896fe5d1dd56345e3ab3541a8034748cbfc411c109c896185a

Observation a0c110ff-ee9b-45a9-8cad-83a1c8769821 · outbound

This paper cites By the same argument as in the proof of Lemma A.7, there exist functions ϕ:R d0 →R d1 andρ:R d1 →R dout such that for everya= (a 1.

Training Transformers for KV Cache Compressibility By the same argument as in the proof of Lemma A.7, there exist functions ϕ:R d0 →R d1 andρ:R d1 →R dout such that for everya= (a 1

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.950355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:01:40.766843Z digest=sha256:3eca651ef5376e305a0f1a23c621e8da14064c1547ac9d353adc986f1faf0eb4

Pith citing papers

Observation 7d34a2e8-94d2-4583-a0d8-ad9398ba8fba · inbound

Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches cites this paper.

Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches Training Transformers for KV Cache Compressibility

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.750886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:35:22.750886Z digest=sha256:1feca56646c76c054f5ad665ae7ae6e14abbd624573e399ed588eb868803d29c

Observation 910c87a8-5fa8-45c2-84b1-25cfb63079ee · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory Training Transformers for KV Cache Compressibility

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:44.516747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:44.516747Z digest=sha256:62deaabba74a13c8c08d64909aff1ae1f33703e7d22be80f3f9817a3ee43bef2