Pith. sign in

Paper Citation Record · LEDGER

Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2401.07851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.07851 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:42:44.390745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:50:04.384621Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e7c8b869-cec5-4ecb-ad09-032e8276c2cf · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:23:49.645933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:a5fb8a848db28b66cb09de091d44a94bbe80f26cb284c2d82d8df292116cdd3c

Observation 3d69fc85-6b0c-442c-b0d4-c8660aa56cb9 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.456272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:fc8a1511039ae004569c4ac45682168121af68e6547bbe00b330ccafb88be3a5

Observation e84f72ca-173d-40ef-ae74-e116df59bb97 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.555244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:2caa07c3b3869fe7b7a26849c12a71319329235ab6d11299b217427bd6265378

Observation d4d370f4-c46f-424b-98f5-db247d902e32 · inbound

Reward-Guided Speculative Decoding for Efficient LLM Reasoning cites this paper.

Reward-Guided Speculative Decoding for Efficient LLM Reasoning Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T20:42:44.390745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:42:44.390745Z digest=sha256:640fea15f02f283ae408225a73d3f6556fbdd9c05922cebaba2a7f12b75c698c

Observation 771287dc-540e-4263-a3ca-95401869bafc · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.795746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.795746Z digest=sha256:b869e1b64658cab2926d2d74d8e31092177a68668f4b78bac61a5ebd977e2a26

Observation a87a8082-5fe0-4a6c-a082-c99ae2dc746a · inbound

Reflection-Window Decoding: Text Generation with Selective Refinement cites this paper.

Reflection-Window Decoding: Text Generation with Selective Refinement Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T04:13:44.424825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:13:44.424825Z digest=sha256:bbc18a716f63032065741f872449a8909e30e8715311a6ab333c39f4a04fcf36

Observation 8f738c5d-bf30-4c14-9d3d-1108f790f33a · inbound

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE cites this paper.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.993285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.993285Z digest=sha256:1ba5bb66e8bcce5b208c295958327bd2564f8370f05e6c1132eed7fd955ef482

Observation 4b6a643e-a163-4897-818a-5763be994129 · inbound

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding cites this paper.

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T11:11:44.653199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:11:44.653199Z digest=sha256:1ed4c74cb731c975000d8d6b04d9b3883a0d02cad8c15d45e8568504b05b7796

Observation 351ee2cc-f9ba-452b-b620-bb0869482f2a · inbound

VeriThinker: Learning to Verify Makes Reasoning Model Efficient cites this paper.

VeriThinker: Learning to Verify Makes Reasoning Model Efficient Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:14.802495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:14.802495Z digest=sha256:c0be273bf7c31922117216a8d791f02f245044110b8186f14f74731af253d2a6

Observation 3aea8502-c3da-49e5-bd95-5429f5c2cbae · inbound

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding cites this paper.

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.749845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:42.749845Z digest=sha256:f40680ae9af5762ec01101eaf022415410bb1ee1e56d16650f39d6a884e3ed3d

Observation f9be1957-6551-4f7e-b129-df13e605b179 · inbound

Consultant Decoding: Yet Another Synergistic Mechanism cites this paper.

Consultant Decoding: Yet Another Synergistic Mechanism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:57.001513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:57.001513Z digest=sha256:af1c438f308dcd98cf046b7de72ef53196774df658358b946926a8a3becd9bf0

Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · inbound

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism cites this paper.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.384667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.384667Z digest=sha256:01dcb26474bf52011db477a7bd3b54656ffcc4a4bff398bb4dfab047d8714581

Observation 03a5e258-919c-4c63-9356-739963217701 · inbound

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models cites this paper.

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:46.418338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:22:46.418338Z digest=sha256:ce0f14bdb11b9de3102aff92f8997efc11c15943caf12f029e3e6f5e6b6a6db8

Observation 98267ba7-8eeb-4aeb-9e78-24ce97bd8e69 · inbound

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference cites this paper.

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:44.088718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:17:44.088718Z digest=sha256:45735ac434a9d599f73c7029279ff7025e0a3eb5d2e62b36d00843b8cf8903d0

Observation 4e49a438-df9f-4712-870a-40c447917794 · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.884383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:402c0ddf629eb263ada86b5585b3bb1b8d9834ce6d1fe6c4ffbd96a27c6e91ab

Observation 6291821a-e0de-4494-8be3-44ba6674a1fc · inbound

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios cites this paper.

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:10:03.112922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T14:09:59.938676Z digest=sha256:d97c26c74c55b40f186220203558cd9b036756ee83887069fcf400a99563d144

Observation bbc09dcd-f2d6-4dd4-8436-93db38143891 · inbound

FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving cites this paper.

FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:59:17.685404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T22:56:19.734262Z digest=sha256:6883719be6d2a74fcc8221d52a23f9c43f67afda1fe74895570d076fae477373

Observation e609ecc1-d828-49f8-8e06-529c24ee3a5a · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.385717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:e9364299487a1261489bf59f14685283a442b18d476ce3b0d7d63102ce9acffa

Observation 80a3d0c8-c0ae-4104-a2ef-b64929f2a1a5 · inbound

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? cites this paper.

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:21:26.630154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T11:08:25.132013Z digest=sha256:99c45882645c18eecc270b698d032143ed4ac7ca711df10fc50843955367c34b

Observation 45b71c1b-b261-4f19-8aba-346a2c817270 · inbound

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? cites this paper.

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:23.953655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:59:11.690911Z digest=sha256:c813b69d2b75e979c70de9e32c5bc6599f2b51006793c7af9b3bc85d0bf32e47

Observation c3f4aeeb-d47d-416f-8446-7e6dc91f13e3 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.281087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:f86d45126c0218ceb000fb6f1f7ad16b07a2e3683d9d3f8285cde2a989845236

Observation 8313ebec-ae47-469f-8d9b-6f9299776a4f · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.451418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:f205e0d7fd79c5480b7fca7196dc0f189aac8bdfc3be694fe2c8d20fe1615801

Observation 86141a99-b4a2-44e1-9db8-182f15452ebe · inbound

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting cites this paper.

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.117035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T21:56:42.264380Z digest=sha256:1f3b246c139edc5c10a4634619dcad5210505e167b6b1687ceeeb0e90b8ff9bb

Observation 8b7018ec-e240-42a5-84ee-d0a9083ec63c · inbound

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding cites this paper.

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.426899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T06:58:07.996335Z digest=sha256:b53c6654fab0fe35d95dadd33b1831a30057ddd36b03e48da078e17399ac9d12

Observation def9a5c9-d4c8-4a89-af31-72992f9fc9d2 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.154232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:18853cce5766f249a52df15ecfb14481abba7e83b7b755b88357c4cea64519b2

Observation eddf3414-193d-46a8-ac71-c15c7f02f5c3 · inbound

SURF: Separation via Unsupervised Remixing Flow cites this paper.

SURF: Separation via Unsupervised Remixing Flow Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:16:52.609778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T05:07:10.235599Z digest=sha256:08c09b3ed1ebfab95dd74d44bc9d66eab84dd254dd276a73def4449ea3c74aa7

Observation 03d59875-df8e-4692-a416-a10eb5d24238 · inbound

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off cites this paper.

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:50:04.386306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:31:27.598322Z digest=sha256:78f902a50070713ab21c965a8c2e4a8da01c1e347277ab24d210537a4561fc6a

Observation 9a768270-878e-4961-9570-8d80e02cc279 · inbound

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding cites this paper.

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:46.570419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:46.570419Z digest=sha256:1c9369a7ce0f7aacd6f112d13340754a37f21bab65df7facb6409a5b8f41d070

Observation 3d44352d-c718-4c95-81c8-019008783cb1 · inbound

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference cites this paper.

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T10:14:14.991489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:14:14.991489Z digest=sha256:5f49cb786fb43d943f596d12c3e167368562097236ede34b9878db661bf1b6e5

Observation 9b1e2b73-e24a-4ae4-9db5-7ddffd24fded · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:13.844047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:13.844047Z digest=sha256:8df700612556f0b02855b02ee9f6d227c960b11e00fb7519ee8e5fef7a1af938

Observation 5389dc48-786e-454d-9406-0d90fd08b364 · inbound

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes cites this paper.

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:06:58.268790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:06:58.268790Z digest=sha256:56ca7706aa32216ae267b904c2e2729846d4f7cc65041dc0dcf6222a69b693f8