Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:54:14.123344Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.04074.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:54:14.123344Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dffa9505-32f4-4a44-8b11-943e79995d5d · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms gpt-oss-120b & gpt-oss-20b Model Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2463839-7179-43f1-8f21-facf893d71e5 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc846ac-aa8f-4dca-b295-7140eb5ea9d3 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7ebaf411-d1b1-49d8-94fa-34c1c5beb7ae · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms LooGLE: Can Long-Context Language Models Understand Long Contexts?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e17b076-902a-442c-9477-b77ee178bbda · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms CommVQ: Commutative Vector Quantization for KV Cache Compression
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5884f49d-38b9-46de-87c7-b7891340b358 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Jamba: A Hybrid Transformer-Mamba Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683bf42f-6190-4a5d-8f23-abce0a8533fd · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Let’s verify step by step
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6448dbdd-70ab-4cc5-993a-64b0028f97d3 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00318c5f-ff3f-4a87-ab9e-71b2e0d5eee2 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c26530b-5a78-49b2-b132-e37d20363a2a · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 28187aa7-6b66-411d-acbf-b8e988321572 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Fast Transformer Decoding: One Write-Head is All You Need
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a05d0bf4-a162-4d4e-af95-648437d2875e · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa47726-d9bf-4943-a713-36c852ab965a · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Efficient streaming language models with attention sinks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2a34392d-b271-4c82-855c-cad61477afb9 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Qwen3 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d86e9c-3cb9-4aaf-ac2d-639523fdbbef · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a97561-1927-4eae-994c-3b2b5b9b3412 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
Reference 1949
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae330aa2-8bc2-43e9-aa41-b3782d315ff2 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
Reference 1971
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9845232d-b613-4b20-a425-e9ae727f475d · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
Reference 1982
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bacb15e-2bbe-45f4-b93c-5198d689821a · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms AIME 2025: American invitational mathematics examination.https://maa
Reference 1989
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4e55d921-80a1-4a16-afc3-37cce8ae555b · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms doi: 10.1007/978-1-4615-3626-0
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68dad4eb-3689-4e02-b479-8c4761700957 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Gemma 3 Technical Report
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca02879d-05da-4900-b1e0-7cf66f6e422c · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008b43d0-b42b-4e71-bc8c-c7cd9ac68b61 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms The Llama 3 Herd of Models
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff032e55-9615-4709-88a9-71175f57a4f8 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fd2b8d-8ee7-44b9-8920-d8d4c513ef0b · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Kitty: Accurate and efficient 2-bit kv cache quantization with dynamic channel-wise precision boost.arXiv preprint arXiv:2511.18643,
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f52db0ef-2c90-4882-9b9f-6a74d1d43053 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Expected attention: Kv cache compression by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c25ae6c-ede3-45bb-985f-35df6b90644b · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Evaluating Large Language Models Trained on Code
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affaf05b-ca25-4a55-ab28-a150aa3d4916 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098c89fd-a716-4d80-b679-63aa5339b356 · outbound
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.