Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:28:04.009153Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 4 inbound Pith citation observations for arXiv:2502.10424.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:28:04.009153Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:18:31.708860Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:49:57.974789Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b5834700-2e79-42ba-b437-8b043f26406b · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0c95a0-b988-40f2-a1c1-9201997f59d9 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Accelerating Large Language Model Decoding with Speculative Sampling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67aba63f-2798-4df1-9f36-1e7f316667f5 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Scaling FP8 training to trillion-token LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a42e937-f4c5-4818-95b8-56574eddf90d · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586eed45-5c2f-4321-9e0c-ee9322a13662 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b679b5-cb6e-4dbc-b430-5e0872a62807 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache KV Prediction for Improved Time to First Token
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31dcdd4-0ac6-4064-8f01-5ef9ccc71d1c · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9949b154-5c0b-4936-8e4c-06b655bb5173 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SqueezeLLM: Dense-and-Sparse Quantization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc780161-7c62-47b6-be59-14fe62cf56ee · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SnapKV: LLM Knows What You are Looking for Before Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa6e5d2-454d-439a-86e6-6b7b7b112699 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7345ea2b-2c4f-431d-9eca-c27c29b848ff · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b4d8c2-04c8-4e6c-92e6-dc57740d77a2 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FP8-LM: Training FP8 Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e1fb73-74f7-40d1-b9c5-c379f6b87548 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad286d0-5105-4022-8ddd-c8bd2ec549bd · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08e0a6d4-7da1-4fa0-8e17-c8c2fd0ae098 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache LLoCO: Learning Long Contexts Offline
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86de0b1e-cb5a-4131-9d4c-b24c85177233 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1038716-a5c5-420a-ab12-29e19053edf9 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SirLLM: Streaming Infinite Retentive LLM
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a313e2-e6e0-47ca-99dd-de52c3fe93f1 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15dcf886-47a0-46b8-9aec-886205754ddd · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Sageatten- tion: Accurate 8-bit attention for plug-and-play inference acceleration
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89c69067-4879-4461-a393-8a3255bd9f70 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Sirius: Contextual Sparsity with Correction for Efficient LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ba0280-3b35-4efc-94af-26e76cb50cbd · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Attention Module’s Inference Workflow The inference of LLMs can be divided into 2 parts: the prefill stage and the decoding stage
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 381e4955-fb7f-41b2-8a31-e7cd894694d5 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9fa495e-ab32-4716-8992-60667af416e6 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache • C4 (Raffel et al., 2020): C4 is a large scale web-crawled language modelling dataset mostly used for pretraining LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8178b162-b3f3-46b6-a433-7affa2923587 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Unresolved cited work
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22feb461-c743-4065-9d5a-f9a8fd9ae6d4 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9058b287-13a9-472f-8692-f8d0b39a65b5 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6ab705-b51c-407f-82c4-17636565bcfb · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Compressive Transformers for Long-Range Sequence Modelling
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335e973b-de6e-4ba3-92bf-c5d1b97644c2 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65a207b7-39fa-4add-ac92-47e2d62c9e33 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd10325-3177-4fc1-ba60-2ea109fdaa68 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08105f62-3cb3-4c3d-94e8-4f4c00371321 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30d4b816-ebab-44fb-b757-ee0633875c68 · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c04e46d-7b85-4a4a-89a6-1e6ba3cd289f · outbound
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Speculative Streaming: Fast LLM Inference without Auxiliary Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13708c35-1129-4200-9f60-9ec760260725 · inbound
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0656267-ad41-4e6f-87a4-232f400ca9a7 · inbound
Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0ab90f20-25c1-40fd-806a-42174a472642 · inbound
Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eae26172-954a-4608-af4f-ba240ae149de · inbound
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.