Pith. sign in

Paper Citation Record · LEDGER

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2508.10395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10395 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:37:15.720135Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c1eef06-70f8-4abe-a5f2-a32206256e35 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.313365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:11.718326Z digest=sha256:aa59305d6626f297772064e49a2dae2d7f1639245988d28d376a01d0b474c593

Observation 7b5c91f9-60f9-4118-a663-eaaae8bdc347 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Longbench: A bilingual, multitask benchmark for long context understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.302365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:11.802675Z digest=sha256:05f20e5b50ae921702576f3e78feb6870b8cd01b1a479c44d53783a03f4a6ab4

Observation 3f524b65-db83-47c7-aff0-08794c3652b9 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:11.866338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:11.866338Z digest=sha256:44da08327980745c81976fcb0943488f5a3bb0f867d169a72aab9dadaabf616b

Observation 2323e10e-2681-401c-858c-2587f9bb0dbb · outbound

This paper cites xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:11.923378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:11.923378Z digest=sha256:defe9ea6032628dde3cbcfa37c5830869d16ecd89843bf2f1675804645106f25

Observation d4076f9b-e933-4659-91f9-57fa21f5e753 · outbound

This paper cites Palu: Compressing kv-cache with low-rank projection.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Palu: Compressing kv-cache with low-rank projection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.283985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:12.013603Z digest=sha256:e92ad392343bad5ad1537d0eb68cda19cf3cc43f89c29178934cc106a0dc0941

Observation afbd76c2-0382-4128-88f0-b7113a4bcdad · outbound

This paper cites Palm: Scaling language modeling with pathways.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Palm: Scaling language modeling with pathways

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.072745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.072745Z digest=sha256:85134a7f41380f75d4b85b0204bd43d9a657c0e6e3ccfffc834a32f47e0b44cf

Observation 8f1ba5c9-83b4-4e4c-90bf-d2082ebd2070 · outbound

This paper cites Training verifiers to solve math word problems, 2021.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Training verifiers to solve math word problems, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.157932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.157932Z digest=sha256:d44b42aa4fb312e30c2127eef2ae20fde1ab4c38dc0e654928c8b60c8b3b27b7

Observation 15832a7a-6a4d-4b8e-ae9c-372c89b42146 · outbound

This paper cites The Llama 3 Herd of Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.244737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.244737Z digest=sha256:203a78251231fd63dcecb160ab32d91c4bde77d14a8edae9110708c16b5dde5a

Observation 0aa150db-c316-4ac7-98f4-70d155a558a1 · outbound

This paper cites The language model evaluation harness, 07 2024.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization The language model evaluation harness, 07 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.334012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.334012Z digest=sha256:dc057da50258a5f3a95b7e3d6d0656269ce1dd4a27565b1a1aa2d376a8e36b8a

Observation c15ce080-4ef0-490b-a6b5-09817e47bc31 · outbound

This paper cites Fast state restoration in llm serving with hcache.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Fast state restoration in llm serving with hcache

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.250581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:12.399294Z digest=sha256:9d6b2d6798a11ec26819dfd1789144de3bcbc0964b203488685a5e934de1659d

Observation 7e72e24a-621d-4c84-b7fc-4b430db5f4b6 · outbound

This paper cites Ai and memory wall.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Ai and memory wall

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.239833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:12.480539Z digest=sha256:0dae4b83b139d3f0223f8c70518e120ef447a1fa52aeea7cddbca1e626a02a5d

Observation e5e2d711-c226-46b9-899b-1bb84e4f5ea4 · outbound

This paper cites Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.575486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.575486Z digest=sha256:7a613c010878c2391b224c594e2d580323afbcae7d297c037bb73fa6a90b9f44

Observation 9855f0ae-3af8-485a-be6a-c2bcf512839b · outbound

This paper cites PolarQuant: Quantizing KV Caches with Polar Transformation.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization PolarQuant: Quantizing KV Caches with Polar Transformation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.660388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.660388Z digest=sha256:758b9c0b8dade1736e65b69cba3105479e4bba4bf69ded538c9b9d8ccb4ed93c

Observation 3968a25b-49c7-4b2c-a48d-fec801aa1f4c · outbound

This paper cites Zipcache: Accurate and efficient kv cache quantization with salient token identification.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Zipcache: Accurate and efficient kv cache quantization with salient token identification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.228837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:12.754333Z digest=sha256:b1d8235b551b2be429dc3bc3a4325cfa7468007b5672f812d408d447b274db80

Observation 631dec7d-3102-4409-a021-98f27c8b7cb3 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.874359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.874359Z digest=sha256:df6371da63a448011f746c770def0a8e8e7f56d62781f4e5632beb284fb47d20

Observation a2d231dd-b670-40fb-aa4a-bff9f51e1fe8 · outbound

This paper cites Checkmate: Breaking the memory wall with optimal tensor rematerialization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Checkmate: Breaking the memory wall with optimal tensor rematerialization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:12.989497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:12.989497Z digest=sha256:eb4fd4f04f86c5738cc5298293b11074e64c0ea0ff50b8cc51a2ce5de380247a

Observation ce27969a-f689-4ccb-9730-79aa815ebc1d · outbound

This paper cites Residual connections encourage iterative inference.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Residual connections encourage iterative inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.205712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.070387Z digest=sha256:f55a71ab2e42c38776d814df1b313b9ec02ff4a4c9e4be8f4cf7ddacf6b98264

Observation 4b9171b1-b443-4307-925a-a6c3eb5e615d · outbound

This paper cites Mistral 7B.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Mistral 7B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.135459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.135459Z digest=sha256:4d4bc65a5a71d351629496785fab169bbeebd38ca9930f39f709757515880c1d

Observation 6c9e2511-d8cd-4035-97bf-314a49baa76e · outbound

This paper cites Mahoney, Sophia Shao, and Amir Gholami.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Mahoney, Sophia Shao, and Amir Gholami

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.195749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.221672Z digest=sha256:58ab5bb8dee959a59ce8994acb051318020b8dea53ca98e8d3bf657e294e802a

Observation 16de2b6d-3f68-4667-ae85-a82b0265a614 · outbound

This paper cites Squeezellm: Dense-and-sparse quantization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Squeezellm: Dense-and-sparse quantization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.186133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.297646Z digest=sha256:d949a3289a7e851a87a1933c354aedddcb3c5459941d279e7552da85517fdb6a

Observation 66d614fa-0a90-4ecb-9a66-70505ce68cc1 · outbound

This paper cites Efficient llm inference with activation checkpointing and hybrid caching.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Efficient llm inference with activation checkpointing and hybrid caching

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.347682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.347682Z digest=sha256:f4152a6e17b7f8a9e7ec6ac7f8ba3a28b4efa72b3dfb3055cf8fd410b2e3f720

Observation aac11420-51eb-4b86-9915-ce5839478833 · outbound

This paper cites Intactkv: Improving large language model quantization by keeping pivot tokens intact.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Intactkv: Improving large language model quantization by keeping pivot tokens intact

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.176321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.424917Z digest=sha256:8079c63fb54855f5093c1ba53f309349e2e70626c195a765402eeeb1c3c4c485

Observation 6ec37ac1-38a1-4b1f-993f-383e9afd8792 · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Kivi: A tuning-free asymmetric 2bit quantization for kv cache

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.510276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.510276Z digest=sha256:7b9b6504ec07067fd30671c58eeaa912b336d25accf0daf1ecd04ceb5eec710a

Observation 5075b5bd-4990-41d5-9345-1bff31bedf0c · outbound

This paper cites Pointer sentinel mixture models, 2016.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Pointer sentinel mixture models, 2016

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.576293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.576293Z digest=sha256:ea67d40d61f26884defc75e0fc3764fb31118f48da6fd0bb94f40e9c71e94202

Observation 2a402916-1310-475a-9143-b456d4ba0e3c · outbound

This paper cites Llama 3.1: https://ai.meta.com/blog/meta-llama-3-1 , 2024.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Llama 3.1: https://ai.meta.com/blog/meta-llama-3-1 , 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.154060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.666970Z digest=sha256:0091c41d19059e36a6c94003962711cc4589101595ce01814ab698d317f5a385

Observation 0a2e3a84-3e2c-44d5-bbdd-27c0cbbb5682 · outbound

This paper cites an unresolved cited work.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:13.746229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:13.746229Z digest=sha256:12386f9a8837366a21ad08444cae780425fd14a2c0c176fa8312e3e54cdf9316

Observation 07d52c83-fe05-43e7-b10e-22b823fdc017 · outbound

This paper cites Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.137087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.829043Z digest=sha256:4c0ef3ad0695ebd122c1c87735e33f9f504d2fddac9d9686088e04bbfaac0537

Observation 1c5359fa-7684-4b7f-928f-63b05c1448c7 · outbound

This paper cites Eigen attention: Attention in low-rank space for kv cache compression.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Eigen attention: Attention in low-rank space for kv cache compression

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.126325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.892816Z digest=sha256:f2e2f112a14dbdc2b19c2d2873dad7eba9e34b734c572b096f3c813be30b461e

Observation fdbb958f-2b72-49ba-abe2-c377589a0e98 · outbound

This paper cites Loki: Low-rank keys for efficient sparse attention.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Loki: Low-rank keys for efficient sparse attention

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.116849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:13.962201Z digest=sha256:1edb9a3833f0019285f7bdb442c0576dedc2f0a4874b367e63148d3e3541402d

Observation 7d51c3f3-3efb-44ae-9b02-a22e32eeed36 · outbound

This paper cites Roformer: Enhanced trans- former with rotary position embedding.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Roformer: Enhanced trans- former with rotary position embedding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.107223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:14.040351Z digest=sha256:cfc6fcc299aa6182a5a49846577491ed20208ae782a0543c06a9ad229e484604

Observation 06966c43-271f-4c7e-8221-fa75e02016c8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Gemini: A Family of Highly Capable Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.106347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.106347Z digest=sha256:cf169a794f7c1c496c354407db033a4530b31877d5b5665b5964fb04da483c79

Observation 136be86f-4dfc-459c-addb-b7adc00bee00 · outbound

This paper cites an unresolved cited work.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:37:16.097187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:14.177445Z digest=sha256:ca7e6f3e04246bc23813d61ab30fd138d05b91dc421296ebbfbc66e6d3982c8b

Observation 7c385cdb-0eda-40c7-b86b-c8145ab0f465 · outbound

This paper cites Quantspec: Self-speculative decoding with hierarchical quantized kv cache.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Quantspec: Self-speculative decoding with hierarchical quantized kv cache

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.087682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:14.251524Z digest=sha256:db3a56f01fdb43a94f34a27803b8aaaa4ac7e34b4b1b2d4e6766859ccabb5e69

Observation 047f1126-b2f5-4fa9-9538-fa1d4a02bbd8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.314482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.314482Z digest=sha256:8d0feb26fa820d6cb9b9756dd8d39ea91023dec47f032f442d4434deaf5aa868

Observation 0ac099aa-ddd4-4725-bf1a-129e143d515e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.389766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.389766Z digest=sha256:ecf2c4d0b4d4eb35187c42c250fc11dc8c88b02686cba1be212890528f30d100

Observation 4a0741e7-31e9-4492-bafa-6d4b2456f1bf · outbound

This paper cites Attention is all you need.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Attention is all you need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.452375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.452375Z digest=sha256:1234ae65e7f9ec98336a6be5ab4bec536cfa51977892cf664c5c85f1c20aba30

Observation 4745bc3d-d048-4169-bed1-831077355f13 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Roofline: an insightful visual performance model for multicore architectures

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.495134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.495134Z digest=sha256:8e4e33e899646432ae09c1616014f96d25e72f329a7382f26475592c4c04280a

Observation 3720f747-8b2a-4343-9bae-c838fb1fd24c · outbound

This paper cites PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.599574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.599574Z digest=sha256:2c2a2159b713f75e6c3672b94803befa0d91e428aa0a268f8099ee61c08cd477

Observation 15ecedfa-918c-4f51-bf87-f9124faf0059 · outbound

This paper cites Efficient streaming language models with attention sinks.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Efficient streaming language models with attention sinks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.064850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:14.615011Z digest=sha256:e619811d31df2ad98f6c1e401eac8f077ac1e206727e5ac546dfeddf5f312136

Observation f5905c12-24a7-4e6a-b390-da8ece273cf7 · outbound

This paper cites On layer normalization in the transformer architecture.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization On layer normalization in the transformer architecture

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.637536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.637536Z digest=sha256:3eaf68c26201ae81b4c32358751e609363b4a629f95a27c0a57eab6570b15f98

Observation cef110ff-40c1-4fcc-b6d4-7d18d26e9899 · outbound

This paper cites Recalkv: Low-rank kv cache compression via head reordering and offline calibration.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Recalkv: Low-rank kv cache compression via head reordering and offline calibration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:14.749127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:14.749127Z digest=sha256:f55e7ee47ace16c9e1fe459c4b6636147bf26e4372e22c10caca388128e41ef4

Observation 56482788-ff5a-4cc4-b718-e7969d227cea · outbound

This paper cites El- attention: Memory efficient lossless attention for generation.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization El- attention: Memory efficient lossless attention for generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.048265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:14.924191Z digest=sha256:630fce740a55bfa0b1b93d04bf6375d8997fb26f784d28fb6182ccc233ed9824

Observation c7653971-b57a-42d5-8ea4-c3692c18ec05 · outbound

This paper cites Qwen3 technical report, 2025.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Qwen3 technical report, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.107048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.107048Z digest=sha256:20a170ad469f88a254f8fcc57915f6305a5533f1d52d192ea7dee33f8f5ed3f3

Observation 45cd5222-97fa-4974-a33c-7dbc48f255f2 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.322633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.322633Z digest=sha256:ec9101d2991a06f75fda4a167abfe70c9bca8727f327b2cfdb3fb0f61ed07eeb

Observation 84d11a4a-1084-4bf8-812f-086dfa615dee · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.475893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.475893Z digest=sha256:309ad9ea81ceec4df99aca552cfa6256d1e2de293ac3a4d6fbdcfe19049de93a

Observation 88e694ca-5ae1-49eb-acb9-af558ae55074 · outbound

This paper cites Root mean square layer normalization.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Root mean square layer normalization

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:37:16.031243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:15.673100Z digest=sha256:270420e7895d7b0c21831029739603799e043fb4d3d5ea5d2f5eac972276f548

Observation 594bfca9-8323-4366-8d46-5c149d425ec0 · outbound

This paper cites LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.716757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.716757Z digest=sha256:64ae3a303ddbf8fc53e06697b83d2225c49a3e143084864e7164a5afd2e7157f

Observation 94d290ba-be83-46d2-a099-fdd937a0a2c2 · outbound

This paper cites an unresolved cited work.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization Unresolved cited work

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T20:37:16.020486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:37:15.720135Z digest=sha256:734f874342744691d076f6bde99c8102fea613b6e3ef39c9dc3c76423c9bdcec

Pith citing papers

No inbound Pith citation observations are available.