Pith. sign in

Paper Citation Record · LEDGER

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2411.13157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13157 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:49:22.119488Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:14:49.396548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T05:52:37.366136Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact4
  • verified fuzzy1
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0abdf2cf-8cde-4e6f-b36d-9eb5b1a2bb38 · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.949062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.949062Z digest=sha256:81332c19f0d0833349febe1b9b3817aed55b8ed58f474971f17babf86058efcc

Observation ee22bae8-de6f-4823-ba50-06861346e704 · outbound

This paper cites Language Models are Few-Shot Learners.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Language Models are Few-Shot Learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.954693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.954693Z digest=sha256:ff92822efdc071511e16de7f7aa5bb95e8a74a8299f6c707920187a79abce037

Observation b2de981b-9f46-46ff-b7ba-f463d6da1a1d · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.959474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.959474Z digest=sha256:3d60ccec005a0dbbfe63fb0adee60721db1c63ce016822dc189af63d92c2110a

Observation be596067-787b-4f63-ae93-cf5097a769d3 · outbound

This paper cites Venieris, and Hongxiang Fan.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Venieris, and Hongxiang Fan

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.963961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.963961Z digest=sha256:e9f61a58f1e574376855380232f2f4fbb2ed14031630c09f14c6bd32f0e670a8

Observation c6975d42-9037-49d9-898a-845f03dc45e9 · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.968003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.968003Z digest=sha256:afab163595ae2825edd4b79641e537ab49cfa8d8ce2d2659d417003a0bcb11fe

Observation 07188d4d-5a22-405b-bb08-22ab47ce7331 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.972377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.972377Z digest=sha256:f8f05fec43bce83c487a2bbbf38d8bc76b13eb844a79bbc157f5568d1d1e6676

Observation f1e8f27e-4c36-4d5b-a8dd-e8e522e1cc1b · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.976799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.976799Z digest=sha256:bba72afd0bd0f6afad5fe6a2089da39e95b2bbae0d3e383526e6b1d8c95c57f6

Observation abe529a7-b683-4686-ab18-64dd79546dc9 · outbound

This paper cites Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.980587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.980587Z digest=sha256:ca9ef46eccac063940f7adab4a428bd968a4dc41bbae07c2aa7e625f6b825a84

Observation 874cc401-3e53-48d5-a40a-e7cdf8612913 · outbound

This paper cites Graph-Structured Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Graph-Structured Speculative Decoding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:49:22.509465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T16:49:21.984449Z digest=sha256:9643856b6cc8200af4c705af16b0d2edc8cc73d58c660c0ad04dd68ce859be15

Observation f9f8d6c6-2148-456b-b209-0d8740e2e53d · outbound

This paper cites Hennessy and David A.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Hennessy and David A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:49:22.678088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T16:49:21.988466Z digest=sha256:1911f45f3c2218d8f1e6426cb9f7cae5d4092363aa050322f018c185bf86f9d9

Observation 0999b3a0-41dd-41e0-a840-f97b6c267f7a · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding The Curious Case of Neural Text Degeneration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.992107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.992107Z digest=sha256:daf469f8a429776c8aac06ce4bdefe63f78929ee6550c1b13dd2f6b6666ce23e

Observation cc0d3953-17a2-4a76-9e09-025ab2b8bae0 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:21.995860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:21.995860Z digest=sha256:c49d1c5503cf802d60aa0dba7c26446b9de85d2f9f965e64f9d6a63771e5c9aa

Observation 7bee0cd3-585c-4d57-9b91-bd87e18bd513 · outbound

This paper cites Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.001779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.001779Z digest=sha256:7479e926c1ae6b8f884db7c83d139046cedcc91f9d2d1e54a85ca3ea30dd1b25

Observation 8f1f3c4e-f69a-4938-9db7-0c6abfbb3748 · outbound

This paper cites Challenges and Applications of Large Language Models.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Challenges and Applications of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.005913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.005913Z digest=sha256:5b61cc13c0a86949db02cfb16ba87e89688c1184a65b6aa2597095459819de0e

Observation 2807b841-f987-422f-8170-466e7ae4faae · outbound

This paper cites S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.009699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.009699Z digest=sha256:70c1fd0e21ab5dba3144d23385f0bd18016f9e19f7a86be141760b36e3cf1808

Observation 32bcd353-1a6e-4f61-9fb1-6c54de0bb76a · outbound

This paper cites Exploring and Improving Drafts in Blockwise Parallel Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Exploring and Improving Drafts in Blockwise Parallel Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.013856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.013856Z digest=sha256:be158ec868d9cffa1a2c07389f2c5c8e688355eb408a56d9a87bb4b959135be3

Observation fc5a947c-d673-479f-923e-0733027509a9 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Fast Inference from Transformers via Speculative Decoding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.018382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.018382Z digest=sha256:d8df67556b8d14f372e63868d0c769e37e38d954660942ab4db176e1f22da9f9

Observation d36b7f12-c9d6-4ba6-b500-208fba847f01 · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.023146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.023146Z digest=sha256:4c98f4ee1841bbf9cf46c9decf8fa951c8ca1af37a2cc4cb56341658e7ee219c

Observation e5325b75-6181-46e6-9373-c4b5ba1b1d63 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.028419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.028419Z digest=sha256:ebc9e4b06cd63267fde3f65c8b8d7c26e20f249ac8614566ffd079d5ad12d2c2

Observation cc622119-7f13-4af3-9e5a-b57751ecf335 · outbound

This paper cites an unresolved cited work.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-08-12T16:49:22.151546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T16:49:22.033038Z digest=sha256:ea5c8dde71a385c22b4c390766f8a4913174a27352ae04579dddab8a12d29a21

Observation be7a49e3-da9d-4777-a1a0-900695a64818 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.036883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.036883Z digest=sha256:138f892c4483cfe2eaa8f6d76e3cc4bc875581fc29f1809990729ee3b462b030

Observation 6ba9773e-0a74-4e05-a19a-0cae9258f561 · outbound

This paper cites Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.040682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.040682Z digest=sha256:dffd9be0d16006a387032e31845899e2d7d16cc057106e09885b79c04d117dbd

Observation 17217c18-239e-4b89-bd91-465afb186577 · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.044631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.044631Z digest=sha256:a16b65184c16372aa31689eb9d734eaa2526abc5ec4e65fbd03448b89ec4fd66

Observation a1614b59-c72a-415e-a2ab-ed70a576fe69 · outbound

This paper cites Online Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Online Speculative Decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.048449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.048449Z digest=sha256:2275e47a8a432f5c6605300f81ebc2bb61e15e07dde09a73aee67a38bd5b283b

Observation 95c9c5a4-d499-4ee5-9549-a12ad577cb98 · outbound

This paper cites EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.052301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.052301Z digest=sha256:883e77447d87833a1474d3f2f876ca4ed92cefe7cc09f5b39515a97de294a003

Observation 7f645844-ee90-45e6-a8e6-24cf1b32dfad · outbound

This paper cites Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.056194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.056194Z digest=sha256:df3b59f213a5609640a462fb41d4c506c4dde43bac9113e8f3b15102e1e435a4

Observation f275e36b-fef8-4079-b1a3-dee89fa75110 · outbound

This paper cites BASS: Batched Attention-optimized Speculative Sampling.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding BASS: Batched Attention-optimized Speculative Sampling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:49:22.317757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T16:49:22.059842Z digest=sha256:5ee6deb517f42724915ba82e6dc91e9b84cd360c8ed8d55c6832841e2ed34005

Observation 2b2d2d00-3927-4f2c-8462-3032f9a881e9 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Fast Transformer Decoding: One Write-Head is All You Need

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.063495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.063495Z digest=sha256:34ac619a6c556f972d60e20ffac987d32e041de3c5ad2918f462b96c2e34d0cc

Observation 16516d35-5d27-4371-9ef7-d1897b44d049 · outbound

This paper cites A Thorough Examination of Decoding Methods in the Era of LLMs.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding A Thorough Examination of Decoding Methods in the Era of LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.067366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.067366Z digest=sha256:04ed0834145441901c350e5388fc8c4482a71ba18cf303193c85189ff07a0c1b

Observation d952aa2c-dc8d-46d5-bae1-42a5ba7a9efd · outbound

This paper cites Blockwise Parallel Decoding for Deep Autoregressive Models.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Blockwise Parallel Decoding for Deep Autoregressive Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.071435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.071435Z digest=sha256:bd752bf38962a754cbc6e52c0bcae45314674c552a8f09b91bb026ae69f49eb8

Observation fdb70399-e86a-418a-90a7-a6a4e4d6bcbd · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.075191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.075191Z digest=sha256:3b32394ead40592b7e1aff8657c7c0834a469aacf96a87932a3341355b66254a

Observation 84b2bb9d-d6d4-41f5-b1e3-78598bc1ce1c · outbound

This paper cites SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.079379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.079379Z digest=sha256:fa5396d74dc95aa374d2d0b3578a5d2be93afa609fc45f01aed9b9b11a4d6e94

Observation b4aa017e-1ba6-4363-958b-efa98f08795a · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding LaMDA: Language Models for Dialog Applications

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.083189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.083189Z digest=sha256:78208c14bdaeaa09789c7f829697fca6c02510731d736da8d825c920b7a8f817

Observation e9416438-faa4-4589-8eb6-1ca6764a495f · outbound

This paper cites OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.091326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.091326Z digest=sha256:b4de3931d6bde6a1fb5721048e5c7792250a4bda2fef97c758c04bd3273d6765

Observation 917008b4-235a-437f-8a8f-eeb94b1ac049 · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.096162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.096162Z digest=sha256:e41094b34c8f5e806c4bddcc34bdc1197917b0d2b767ac26b2b4b93273283db6

Observation df77fa73-b2b9-41bc-9743-e0bd889b5664 · outbound

This paper cites Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:49:22.207486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T16:49:22.100106Z digest=sha256:fb8721e90cb1f77c33f92b05e604ab405efe4a29523da79441425b2c83fc6b35

Observation 6f858e5b-7290-4b3f-a779-42811aa61ce0 · outbound

This paper cites Learning Harmonized Representations for Speculative Sampling.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Learning Harmonized Representations for Speculative Sampling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.103844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.103844Z digest=sha256:0af6094b3f715d0f8d4b437c0e21ac5a90b2ec587c43f4d0a7ca401bdb5d506b

Observation 65acae3d-50cb-4ef2-8b7a-53141b7509bb · outbound

This paper cites Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.107629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.107629Z digest=sha256:f0c40a34e7f7d63a528cd0e31b9abcb003805c6a84c0f807a1191cd99e27f417

Observation eaa754bb-750c-4a9c-b523-204b3f353780 · outbound

This paper cites ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.111443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.111443Z digest=sha256:2d6ceaf4911050f6d208d6b3761b4662386eb951f2480b5711b0d505b99d6a3a

Observation 95b8f6fc-a444-400e-b818-03f87c1caaa1 · outbound

This paper cites URL: " 'urlintro :=.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding URL: " 'urlintro :=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.115235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.115235Z digest=sha256:2f7dddf4f2fa9f1a332cf90cb3fdf68623b2f9669d61ecb890d4dcdf1200087d

Observation 41d1e8b3-c7e6-4157-870d-e57bf8da292a · outbound

This paper cites write newline.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding write newline

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.119488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.119488Z digest=sha256:a3d75083493a49e73a735cc44300a1e958e4c0bfa2d62d899d1b1dfd9edde240

Pith citing papers

Observation b8429e6c-757e-45cf-b217-b51061460cec · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.369603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:6c5c8ed3cbeb4a718dd233ab564af9faa82349075ef5284d50b10bbd7bd648fc

Observation 1e4581c2-f0a9-4a66-ac11-d89c5b5ce6e1 · inbound

MineDraft: A Framework for Batch Parallel Speculative Decoding cites this paper.

MineDraft: A Framework for Batch Parallel Speculative Decoding Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:14:49.396548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:14:49.396548Z digest=sha256:4193ffb38176244f02c2a1328a98a2c9a13500c2834e2b6fcb3d8813f3afb2e1