Pith. sign in

Paper Citation Record · LEDGER

Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2401.07851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.07851 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:11:44.653199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:50:04.384621Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e7c8b869-cec5-4ecb-ad09-032e8276c2cf · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:23:49.645933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:b3792a3f310b222462f0a3c011d11a17cf86474b7007166289b269d0e4b5fa2d

Observation 3d69fc85-6b0c-442c-b0d4-c8660aa56cb9 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.456272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:f1f67b3084b320594fd20e92e340345cda39aae74895bb86e2d274da17b86e4a

Observation e84f72ca-173d-40ef-ae74-e116df59bb97 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.555244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:d30909727e42d25a478f18c7c33bf047b4378aaaf6804c39f7cbe1764211cecf

Observation 4b6a643e-a163-4897-818a-5763be994129 · inbound

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding cites this paper.

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T11:11:44.653199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:11:44.653199Z digest=sha256:32e3c3cbc66f5d1fa4cb5434d2e11b86ff7faefab292d6c2a87f686aa323fbf0

Observation 351ee2cc-f9ba-452b-b620-bb0869482f2a · inbound

VeriThinker: Learning to Verify Makes Reasoning Model Efficient cites this paper.

VeriThinker: Learning to Verify Makes Reasoning Model Efficient Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:14.802495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:14.802495Z digest=sha256:56ac387da31aa2130a55ba11337bebb6c31f7917238c74d760ee93a6563e1b7e

Observation 3aea8502-c3da-49e5-bd95-5429f5c2cbae · inbound

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding cites this paper.

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.749845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:42.749845Z digest=sha256:89007eda71c1cfab94f8ca79e40d21dc8a420e3b7de781a66d68cc25532fca54

Observation f9be1957-6551-4f7e-b129-df13e605b179 · inbound

Consultant Decoding: Yet Another Synergistic Mechanism cites this paper.

Consultant Decoding: Yet Another Synergistic Mechanism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:57.001513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:57.001513Z digest=sha256:d99bb0d0bee55157c955110e39d38aa6d9cb524db3d235a501597490c6b75286

Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · inbound

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism cites this paper.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.384667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.384667Z digest=sha256:1b2da0d983384704651370fa56abf9a8e8c9e15569d0d4d14a6d6b92fb4b7c74

Observation 03a5e258-919c-4c63-9356-739963217701 · inbound

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models cites this paper.

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:46.418338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:22:46.418338Z digest=sha256:391f98a97320601688e0c600617485ba582d255835bff72b0807d1e84ef1ed7f

Observation 98267ba7-8eeb-4aeb-9e78-24ce97bd8e69 · inbound

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference cites this paper.

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:44.088718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:17:44.088718Z digest=sha256:f2899a02275643bcc89080c44387cf1dc35a4b820ae9767917ff3b8965854291

Observation 4e49a438-df9f-4712-870a-40c447917794 · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.884383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:72076b3a5f375121f35b3ee40b5dc7db73f2c29b2667317e57cb6168b1b6708b

Observation 6291821a-e0de-4494-8be3-44ba6674a1fc · inbound

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios cites this paper.

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:10:03.112922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:09:59.938676Z digest=sha256:7542904128fb62c79d7ecf5db8951eb2953056cbec0c04b0b741da45f8a6c112

Observation bbc09dcd-f2d6-4dd4-8436-93db38143891 · inbound

FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving cites this paper.

FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:59:17.685404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T22:56:19.734262Z digest=sha256:41aa5cce4c00ab8ba2b36084ba879b534503d870a61a2597c583141788304080

Observation e609ecc1-d828-49f8-8e06-529c24ee3a5a · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.385717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:13d1218f7d6c1eabad9abd2ef3c107c1f3ce1e8f48abc758caebed601c95ddad

Observation 80a3d0c8-c0ae-4104-a2ef-b64929f2a1a5 · inbound

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? cites this paper.

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:21:26.630154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T11:08:25.132013Z digest=sha256:48ee2587622ad1faab2550a1ce8fe3b39cfb729ec76eb66ed2ab28885daba802

Observation 45b71c1b-b261-4f19-8aba-346a2c817270 · inbound

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? cites this paper.

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:23.953655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:59:11.690911Z digest=sha256:0cb2664a1cf17b48bf7eb580d06539f1315d2e63dc7be1f75b902d67e2065e90

Observation c3f4aeeb-d47d-416f-8446-7e6dc91f13e3 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.281087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:d6470a798aaebd743294e4675d8a04e649a18f825bba20853b60154df51dfc1d

Observation 8313ebec-ae47-469f-8d9b-6f9299776a4f · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.451418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:fa3f4331f474907c46829411802cc09fc9547de218faac3a8bac4ac1f2bbb9b9

Observation 86141a99-b4a2-44e1-9db8-182f15452ebe · inbound

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting cites this paper.

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.117035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T21:56:42.264380Z digest=sha256:03f6b98f29a55fa549fae1441e32a801c38c0122dff437a18ba89a82b2ea30d9

Observation 8b7018ec-e240-42a5-84ee-d0a9083ec63c · inbound

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding cites this paper.

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.426899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T06:58:07.996335Z digest=sha256:4f86f0bfba114cb2039cbb70d3d13facacf6266bb600d65de5280902a48b8e09

Observation def9a5c9-d4c8-4a89-af31-72992f9fc9d2 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.154232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:f349bdbd3b767690c7664a5ee32f11bec434de2fef7eec50f9bdeb92ed00b40c

Observation eddf3414-193d-46a8-ac71-c15c7f02f5c3 · inbound

SURF: Separation via Unsupervised Remixing Flow cites this paper.

SURF: Separation via Unsupervised Remixing Flow Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:16:52.609778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T05:07:10.235599Z digest=sha256:92908bd74cff3349883b63971372e66e9dfee28fa5af31ddfb20d25df0ac09b0

Observation 03d59875-df8e-4692-a416-a10eb5d24238 · inbound

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off cites this paper.

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:50:04.386306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T22:31:27.598322Z digest=sha256:b3c49b15758dd977bb1ce6acdaeefb7bfa1b08be3dee981ce7c3e9520e1bd78c

Observation 9a768270-878e-4961-9570-8d80e02cc279 · inbound

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding cites this paper.

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:46.570419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:46.570419Z digest=sha256:46418ec9f406f77baf96bbd92ae9fadf8b1aed54b4b73a260175b2804cf47176

Observation 3d44352d-c718-4c95-81c8-019008783cb1 · inbound

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference cites this paper.

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T10:14:14.991489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:14:14.991489Z digest=sha256:2dea8d93863527ea7219fcc5c0d2f4ec9ab4aa72bea4c5a3a6e50d7324201d48

Observation 9b1e2b73-e24a-4ae4-9db5-7ddffd24fded · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:13.844047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:13.844047Z digest=sha256:a4b685ee01080a87e7ffc6fb40272b97d9493039ccae7ce1c15b9ee6baf90c2d

Observation 5389dc48-786e-454d-9406-0d90fd08b364 · inbound

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes cites this paper.

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:06:58.268790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:06:58.268790Z digest=sha256:4fc70dc11b799ab9f29ef09e58b3b950011011c5adb8cc1ad6686e622f598181