Pith. sign in

Paper Citation Record · LEDGER

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache

As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 4 inbound Pith citation observations for arXiv:2502.10424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10424 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:28:04.009153Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:18:31.708860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.974789Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5834700-2e79-42ba-b437-8b043f26406b · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.907233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.907233Z digest=sha256:a9199a68348526671ead215aa7607226545b043c5b26267fc3dd6427be6533bd

Observation 2f0c95a0-b988-40f2-a1c1-9201997f59d9 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Accelerating Large Language Model Decoding with Speculative Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.911019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.911019Z digest=sha256:dad3839671731531019bb02dfe05bb7c723c209b81928d4f32da2ccc8f25b728

Observation 67aba63f-2798-4df1-9f36-1e7f316667f5 · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Scaling FP8 training to trillion-token LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.922252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.922252Z digest=sha256:378466534ee7688b296b02800571a25c3e0110c29fd50d212083d2937fb17ee3

Observation 7a42e937-f4c5-4818-95b8-56574eddf90d · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.925945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.925945Z digest=sha256:5b967201fd62cb66dc7838e7bdc65a741520e617f125ece1446aee63e6a72cc2

Observation 586eed45-5c2f-4321-9e0c-ee9322a13662 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.929459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.929459Z digest=sha256:af1d6baaab456b99573c97ffaa78e4a9a9d5c626a58843200e6a52b9bdfb59a8

Observation 91b679b5-cb6e-4dbc-b430-5e0872a62807 · outbound

This paper cites KV Prediction for Improved Time to First Token.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache KV Prediction for Improved Time to First Token

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.932733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.932733Z digest=sha256:1a93d6eeee82b6967b35eda32e2b3f59817e9772c600f34104c6909029078628

Observation d31dcdd4-0ac6-4064-8f01-5ef9ccc71d1c · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.935822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.935822Z digest=sha256:2ba46f9a6ab96505cae050b6c10ecbae07593138425a174a79e264b3cbb168af

Observation 9949b154-5c0b-4936-8e4c-06b655bb5173 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.938479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.938479Z digest=sha256:bceeddd6c750edc1a0bc481aa0132fe1c48f66ea78dac799aa154eae92bcfacb

Observation fc780161-7c62-47b6-be59-14fe62cf56ee · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SnapKV: LLM Knows What You are Looking for Before Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.941074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.941074Z digest=sha256:e6da5601734c4385a0cc95246d962717adf48759ac55cf46bf44a1681cb1c702

Observation 5fa6e5d2-454d-439a-86e6-6b7b7b112699 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.943737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.943737Z digest=sha256:855807a4e8fa4d012046d8a1504abd6e8d1f5a095268c243703a34dbcc60237c

Observation 7345ea2b-2c4f-431d-9eca-c27c29b848ff · outbound

This paper cites Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.953718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.953718Z digest=sha256:cc5497af70b7adb2b2f06e9e61a5551ad8c473fd6331c00a83bc36382ebb80d5

Observation 15b4d8c2-04c8-4e6c-92e6-dc57740d77a2 · outbound

This paper cites FP8-LM: Training FP8 Large Language Models.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FP8-LM: Training FP8 Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.957098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.957098Z digest=sha256:c45dbe2d0ce5a88f360d536a06faa978b9d3690a76d02d443499f7c459687702

Observation 35e1fb73-74f7-40d1-b9c5-c379f6b87548 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.967798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.967798Z digest=sha256:18d7c50b1fd415fdbf2e659e7dfc29fda93989323b68b04fb3831a3eee29bae9

Observation 9ad286d0-5105-4022-8ddd-c8bd2ec549bd · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.971156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.971156Z digest=sha256:8827f97cf494c23c1b89ea634091cb160c51072f1d25d30101c71c97ba37c4d3

Observation 08e0a6d4-7da1-4fa0-8e17-c8c2fd0ae098 · outbound

This paper cites LLoCO: Learning Long Contexts Offline.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache LLoCO: Learning Long Contexts Offline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.974790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.974790Z digest=sha256:8dc49a35bc1ba98b07414a19c5818b282b2f6952308a0d064c1614671aee978e

Observation 86de0b1e-cb5a-4131-9d4c-b24c85177233 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.978380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.978380Z digest=sha256:a71e60f04ba52894b12dcbe35031f0960dcf49dc443f1a6bd7074786570bba9c

Observation a1038716-a5c5-420a-ab12-29e19053edf9 · outbound

This paper cites SirLLM: Streaming Infinite Retentive LLM.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SirLLM: Streaming Infinite Retentive LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.985750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.985750Z digest=sha256:a7acff513d6ca9af3654dc0f19022e1497cf9385dbc0cd24c738f9c3643bbfa4

Observation d1a313e2-e6e0-47ca-99dd-de52c3fe93f1 · outbound

This paper cites HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.989883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.989883Z digest=sha256:c0abf4d7b13e4f8ed48cded3239034e9aec0d972dd34542a6a57ca5ea8e8de60

Observation 15dcf886-47a0-46b8-9aec-886205754ddd · outbound

This paper cites Sageatten- tion: Accurate 8-bit attention for plug-and-play inference acceleration.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Sageatten- tion: Accurate 8-bit attention for plug-and-play inference acceleration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.993674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.993674Z digest=sha256:46d02cde349305436b0834133bd0e223b3f97d1d9280ec885dcd81e494d07e8f

Observation 89c69067-4879-4461-a393-8a3255bd9f70 · outbound

This paper cites Sirius: Contextual Sparsity with Correction for Efficient LLMs.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Sirius: Contextual Sparsity with Correction for Efficient LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.997075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.997075Z digest=sha256:01ed5319854c2461f98afb928ccf585d22bb1591fc6cda8e9326ab58e3a6a3ae

Observation 45ba0280-3b35-4efc-94af-26e76cb50cbd · outbound

This paper cites Attention Module’s Inference Workflow The inference of LLMs can be divided into 2 parts: the prefill stage and the decoding stage.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Attention Module’s Inference Workflow The inference of LLMs can be divided into 2 parts: the prefill stage and the decoding stage

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:28:04.550039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T04:28:04.000604Z digest=sha256:4a6a4d50a28c434c6d3cf411af5495904ec9ea04acfedcdf85f7e5176ceee278

Observation 381e4955-fb7f-41b2-8a31-e7cd894694d5 · outbound

This paper cites an unresolved cited work.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:28:04.538751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T04:28:04.003607Z digest=sha256:2361aa9c46c5d7c0602b5196a2fae00fe495617a85e79f05eaee2a4d9ba3a338

Observation a9fa495e-ab32-4716-8992-60667af416e6 · outbound

This paper cites • C4 (Raffel et al., 2020): C4 is a large scale web-crawled language modelling dataset mostly used for pretraining LLMs.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache • C4 (Raffel et al., 2020): C4 is a large scale web-crawled language modelling dataset mostly used for pretraining LLMs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:28:04.518602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T04:28:04.009153Z digest=sha256:10d2a7659c7d36fee5bcf14b6785acc4eec046e03b97e1f1a4ab50c56034ba7e

Observation 8178b162-b3f3-46b6-a433-7affa2923587 · outbound

This paper cites an unresolved cited work.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Unresolved cited work

Reference 128

Resolution
unresolved
raw_fallback, observed 2026-08-09T04:28:04.528773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T04:28:04.006424Z digest=sha256:fa31db2d3e01bed3ecffc4bcaa4c68ac341414ee4cceb71a19f84d19e871a2ad

Observation 22feb461-c743-4065-9d5a-f9a8fd9ae6d4 · outbound

This paper cites COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.982110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.982110Z digest=sha256:0a5ef4fb40ca38c242b5f0e97b7750fd254aaa8947a922d49dd993dd32dc1dbb

Observation 9058b287-13a9-472f-8692-f8d0b39a65b5 · outbound

This paper cites Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.950468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.950468Z digest=sha256:9cc40ae11ce9a3c536dff6a19b7289978113d15f281ff646d1201ee4aa89db28

Observation 2e6ab705-b51c-407f-82c4-17636565bcfb · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Compressive Transformers for Long-Range Sequence Modelling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.960513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.960513Z digest=sha256:8d38cded3f87e088e12035dd4a0c55e8d280aebd6fc9c5d996fa36d7e2c8eb3f

Observation 335e973b-de6e-4ba3-92bf-c5d1b97644c2 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.964126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.964126Z digest=sha256:c9966174128a1250f4221b5feb9c4c78e2b90e8bb8ea6689230fbc9a2c829f5c

Observation 65a207b7-39fa-4add-ac92-47e2d62c9e33 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.946924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.946924Z digest=sha256:f171693047c30b83f9df607baa6de2b8e5a99cd7ce5355666dac46b4792556c5

Observation 7bd10325-3177-4fc1-ba60-2ea109fdaa68 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.918605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.918605Z digest=sha256:f981889a783374b070632123ac0822be6a3e8a9067e39d951efc13422d008499

Observation 08105f62-3cb3-4c3d-94e8-4f4c00371321 · outbound

This paper cites INT-FlashAttention: Enabling Flash Attention for INT8 Quantization.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache INT-FlashAttention: Enabling Flash Attention for INT8 Quantization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.914857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.914857Z digest=sha256:eef10915a44c52d2047a79b152c06fdd71dc321081c6c63f5b20c59dde90aeac

Observation 30d4b816-ebab-44fb-b757-ee0633875c68 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.903083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.903083Z digest=sha256:f0b4d07138fccc074f0b7695e72ae080a469c2922b9fe266ead230d751f10437

Observation 4c04e46d-7b85-4a4a-89a6-1e6ba3cd289f · outbound

This paper cites Speculative Streaming: Fast LLM Inference without Auxiliary Models.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.898221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.898221Z digest=sha256:a35f6b625d4691f796e26105a201365b189213ceab4659916a49a4907b834021

Pith citing papers

Observation 13708c35-1129-4200-9f60-9ec760260725 · inbound

Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design cites this paper.

Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:31.708860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:18:31.708860Z digest=sha256:7c8f8474f4c2264b7c37512f277c6e0c1455296d2069aa5d2e403fdafe574160

Observation c0656267-ad41-4e6f-87a4-232f400ca9a7 · inbound

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding cites this paper.

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:25:49.796335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T16:19:54.910430Z digest=sha256:62f643f6e6a1af4d87a950d1dec89d8167e850cf3a99279c588ac50614eac218

Observation 0ab90f20-25c1-40fd-806a-42174a472642 · inbound

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding cites this paper.

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.977064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:11:58.636939Z digest=sha256:5b9d17fd25718efa62b4ef4a0e0a9221fcc7b02eea4ab1ac587162fe0a2bac90

Observation eae26172-954a-4608-af4f-ba240ae149de · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:33.435680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:33.435680Z digest=sha256:e4d285002dd4bef610596976b4acf7fdace7f7c6d173a208d5341f0931507d42