Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 91 inbound Pith citation observations for arXiv:2401.18079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:01:20.089202Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 49358396-5a24-42bf-a26e-8110a97c2a08 · inbound
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f0ce2c05-fd48-4cd7-8700-7480987707f1 · inbound
SGLang: Efficient Execution of Structured Language Model Programs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1d83d1f5-4133-4531-ba96-1b93e15fbeb0 · inbound
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8d622d52-b7d1-44e9-8a56-00fc999812ca · inbound
A Survey on Efficient Inference for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 219
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b6d19c84-6bc0-4145-8471-c17e03e40217 · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b8ca3e16-9a01-4171-917a-ef22f8c98e96 · inbound
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e31e43a2-f16c-464b-a096-57ffd31fb592 · inbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6070763a-9f62-4b58-96cd-0a12f6c51a35 · inbound
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2960792d-8304-48a2-8529-904342be526d · inbound
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf2ff58-02f5-4c4c-91b4-2c10cc73130d · inbound
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b52e2c3-d4b1-4afa-9cdf-4999171ba5d8 · inbound
Scaling New Frontiers: Insights into Large Recommendation Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18c02e2-eb94-4759-aa0c-59bbb2876a97 · inbound
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f659a8-005c-48d9-b736-fb859617fdb4 · inbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad012494-5137-4507-bdda-9643a54b41c4 · inbound
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a4b25d-8ab6-4d48-a875-382b2a14a53b · inbound
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3535e96e-a105-4684-b7c5-8b957cf300b8 · inbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aab0856-dddc-441b-bce0-76553e2988e1 · inbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcde423f-c079-4156-aff8-96a24c3c45aa · inbound
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66bd7130-778b-4b9b-b5f9-bb1cb30f5782 · inbound
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd871bcf-fddd-495a-a514-daf8f60c2f4b · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e256fefb-690f-4ec1-9063-af310016def5 · inbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc330e45-6dd7-4f4d-b307-5b56cc54e84d · inbound
Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75a3079e-0cc1-4050-8109-c0aa5f96f701 · inbound
KVDirect: Distributed Disaggregated LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98dbfdb-7b16-4efd-8aca-1176e80b66db · inbound
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bef686-892a-42fb-bace-ec0f75292a9d · inbound
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 1978
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63704b6d-d0da-4d2f-9966-e36b374e5998 · inbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e95d2d7-7b5f-448a-98e9-7cc2b1147221 · inbound
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d86c255-4a62-48c3-adae-804db6dd5993 · inbound
PolarQuant: Quantizing KV Caches with Polar Transformation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457f745f-79a5-4965-b2cb-17a66955aeb7 · inbound
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe47305-c1fb-4cdd-a013-2f6bdd13060f · inbound
Systematic Outliers in Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecda896b-6e75-443b-8092-fc93d81d32bc · inbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dbfab0e-01dd-4f32-8e8e-538cd0110b33 · inbound
Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238d4a53-ec19-4a44-95f3-f91b2843baa0 · inbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e846879e-f25e-4d65-aa78-a5774b046e17 · inbound
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6d9095-11a4-4ee2-aac3-bd77a67b6b77 · inbound
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e648adde-0231-4344-914a-9972409ac979 · inbound
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc47a00-7d92-4512-be9c-e6aea921184b · inbound
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070ec353-da31-46bf-9097-ec604a14edf7 · inbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a833d507-3149-4d46-84b6-22eb6fd5b50f · inbound
FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 626ad239-cb82-4223-92e2-612a9d4ef34a · inbound
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061c7b6a-8acf-416f-8cd8-aeb6b27e50cf · inbound
Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b846c7a3-8c26-4820-9f30-fd34b7596dbf · inbound
Hardware-Efficient Attention for Fast Decoding KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c23bce-5875-44f6-92fb-259f5d75ee64 · inbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec87567-bce7-4b01-9448-54d41b907068 · inbound
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683195ac-343c-4526-8a60-3fb774b83e18 · inbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f7e5ad-7493-4c4b-bf6d-c01e6c2334aa · inbound
Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c2c40e-455f-493a-aec0-348b59a97461 · inbound
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2baeb153-49e3-4b37-908b-b498371ecb17 · inbound
CommVQ: Commutative Vector Quantization for KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 1984
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c85dd1-eaf2-46ec-b9ce-a79b014b1e82 · inbound
Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b5b577-79bb-4a38-8752-585dbd12bc91 · inbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc9ce03-1f66-45f5-a56c-51d090804ebd · inbound
CaliDrop: KV Cache Compression with Calibration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78e0d83-6f47-466f-8038-3a25652de897 · inbound
OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30df9cd3-da21-4b66-94cc-6e529e1f16c4 · inbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692dbdfc-6c73-4425-9967-499b6355cf2b · inbound
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0ab0ff-6081-457f-8c89-c47ef52b9221 · inbound
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983a32bc-107b-44c9-a3e0-41e1c4621837 · inbound
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 15760c83-790d-4e7a-b65e-b4d4e8dcae70 · inbound
Token Sample Complexity of Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 1963
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e2c0b5-458f-4baf-b766-d3d418fc5841 · inbound
SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d4f3d5-68a8-4f87-9c2e-6f24aedf0a2d · inbound
Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cde0251b-0d4b-4c31-bb72-819a7691390c · inbound
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c7fc36aa-0ad4-44f9-93ab-ccb34322ac62 · inbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d57de22e-4e01-405f-8866-348250b34b03 · inbound
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0cd80202-ee07-453b-b303-5406d11e96aa · inbound
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6413a979-1151-4d0a-b089-2bb3db893e42 · inbound
VORT: Adaptive Power-Law Memory for NLP Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e94c01d3-5644-4303-8353-e0656cc4ef32 · inbound
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f06e10bb-9277-48af-bb23-f84e9d816afa · inbound
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf1e2c40-e430-4817-aeb3-4d4d392feff0 · inbound
Runtime-Certified Bounded-Error Quantized Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4bfb4e11-c9c5-404c-906c-7a5c0666f1d5 · inbound
Adaptive Mass-Segmented KV Compression for Long-Context Reasoning KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1012cfb-6667-488f-bf37-162f4e42edaf · inbound
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4e72868-c583-4321-a84e-da0be5a57f9e · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c5b01e8-7ed2-4d15-b55c-7efbd6e447c3 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f98172-5a50-4385-9606-d3916ec2c9c9 · inbound
STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d00ab5b-bcb2-46ef-b136-993c3c1a8dbf · inbound
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 19ab3f4d-16ff-4062-b6f0-4af4d02abc83 · inbound
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2bd10106-03ec-442b-b763-c59f9b220815 · inbound
GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e5f9032-d7cd-4df2-aa58-ef1b88649739 · inbound
Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318bfd26-f9f5-4a53-a10a-d232fb84b2f9 · inbound
Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 21137c33-95f5-443c-af0c-65450a67d571 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 079a02e3-8f49-4470-85f5-7c3e64a2be09 · inbound
Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370400af-5b24-4d45-9b3e-a0fe5069a658 · inbound
Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c399909b-00ee-449b-aa0e-cdb766908015 · inbound
Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5553b7-0e5a-4377-b8c2-422f584943ce · inbound
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27aab2f4-0d4c-4593-9b3e-55914237ada8 · inbound
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac5d3aa-ed7d-4947-bbca-b783de3c136c · inbound
Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 270ca459-9ac7-4564-81ca-0a1b40cfb920 · inbound
A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c113f63-5339-43b9-95e9-a9b797220e11 · inbound
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c500bb-4bef-4e45-af9e-d8f47a319e9c · inbound
SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d9282e2-7c49-456f-967f-e3304a2e1eac · inbound
SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01ebdca-cafa-45f6-b606-579db731137f · inbound
Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e887a3ac-789a-4f76-9b4f-c99892a87403 · inbound
Hidden Language Consistency Phenomena in Reasoning LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 226
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8271b02-afab-4e8d-b0b4-d7e3fd4ce051 · inbound
When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.