Pith. sign in

Paper Citation Record · LEDGER

Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2309.08168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.08168 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:05:26.239816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:33:24.532038Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2ad079ad-a302-47c2-ac51-8669e7d11c62 · inbound

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty cites this paper.

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:15:49.431096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-15T00:15:49.303458Z digest=sha256:41edade90bf1dd66de676087e9416c9a405b34d28148486bd8e5e002042f4860

Observation 3e5b5439-561a-4942-8da0-b1d09ae336e2 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.348905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:33e7b3d78ac9ef832d2701d334f212cd39a1a32df63261bd03e502312152c63e

Observation 7d6eb76b-9896-49f6-ae4f-52a110ce636b · inbound

FastDraft: How to Train Your Draft cites this paper.

FastDraft: How to Train Your Draft Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:05:26.239816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:05:26.239816Z digest=sha256:a83f89af6ea66aa1c19ce2351fda787a1b1c540cad70eb668487e506a589b663

Observation e0e5002c-af63-483c-9901-77ee3353e1d0 · inbound

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration cites this paper.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.288948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.288948Z digest=sha256:bf3069d47139b17ebe8fd919d69fea9f757f213a76fd20f92a4af67ac02baa6a

Observation 6f8e579c-8b50-4ff5-8380-768aa5ef3a4d · inbound

PLD+: Accelerating LLM inference by leveraging Language Model Artifacts cites this paper.

PLD+: Accelerating LLM inference by leveraging Language Model Artifacts Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:26:59.631815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:26:59.631815Z digest=sha256:d6b497949d67e19004436fa4f35ac546ff9998840479457dbdae8676f388306e

Observation 368bf258-9b1c-4e01-a035-843ea2b7ba09 · inbound

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree cites this paper.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.788351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.788351Z digest=sha256:6acb36c138b12586a3ee67c1b566e5753fc052dac4f3bdb6081eef12bb5ab7d2

Observation 8058edb3-d53b-41ee-a168-34bf101eebb0 · inbound

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels cites this paper.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.584734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.584734Z digest=sha256:96b683b0043f4ba8662b29e66687d858426b62f67e15af16c4e0d2eb8394ddb2

Observation 391f8be7-67b5-437c-9e97-5b1bb0363d1f · inbound

A Survey of Early Exit Deep Neural Networks in NLP cites this paper.

A Survey of Early Exit Deep Neural Networks in NLP Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:46.380797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:46.380797Z digest=sha256:5bc58f44eab7c5dc64d5c2cd185ab99cc053f7d3b66f9de1210e67aa2ba65541

Observation c801eed3-7e36-479b-94ce-f91b886d7a0c · inbound

Reward-Guided Speculative Decoding for Efficient LLM Reasoning cites this paper.

Reward-Guided Speculative Decoding for Efficient LLM Reasoning Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T20:42:44.456802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:42:44.456802Z digest=sha256:53d6af3152d68b25687c41cdf7f1e40d55324a2f903fab95faf3ada8a0feda2c

Observation fec84020-9065-4149-b785-b80c0267ae27 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.819360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.819360Z digest=sha256:e8010bbb2756a04a924c0f3b20967f204f2cb697a26b4c46178d869e827ec119

Observation d6ce2fff-505c-46e5-915d-2f6178d92b1e · inbound

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding cites this paper.

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T11:11:44.668944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:11:44.668944Z digest=sha256:02e9366da6f0a1c27ba373e3c5ba1c469903c03a30cdafa551ebc366ec78ccc5

Observation 2bdbaa1d-2463-486a-abd3-f3156652c70c · inbound

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding cites this paper.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.702035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.702035Z digest=sha256:40a427a22fecc2380f438a519aa424c7ed42d0ad1e733bcd6b21b3a89eb55950

Observation 4d2f5f57-d5cf-443e-aef7-1e292a681079 · inbound

Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs cites this paper.

Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T15:58:16.145109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:58:16.145109Z digest=sha256:519d9f366214919e0d248df49ae660fb35d13d535478d389fc44a1fdc7b1a732

Observation bf776ed2-75d7-434b-beea-8e9b44efacc9 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.505731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:66077d6053829dd8018cb1391d3a33bed3e4452a9ebf5466048208ad996fb3c2

Observation 8d8b645a-6bcf-49ac-8390-3c2a9bc12200 · inbound

DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution cites this paper.

DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:33:24.533399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T12:27:23.866704Z digest=sha256:19a77ff5079e1103ea5ce1e19490e4251f91b716681357995964352f655f142a

Observation aaec6da3-9664-4a7f-91cb-14a99a938cc0 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.638656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:179d1fb0dfc4f6589e2524a3d17cf49bd458e3862f07cd2beaef05e5bf56c2c1

Observation b32bde44-6733-4882-bb56-47e608902053 · inbound

Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware cites this paper.

Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:33:28.405888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:33:28.405888Z digest=sha256:9ed5bd25870442fe8009f4e652bef9b90ffaf036751b8c3a192b29f34de72729

Observation 1c260995-aa48-46a3-a6fd-da6ae9e19d60 · inbound

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding cites this paper.

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T10:55:32.071187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:55:32.071187Z digest=sha256:80100ececc0313d2c2fedb0b5104f9dad999dd76ac58305c55f497d99f610adf