Pith. sign in

Paper Citation Record · LEDGER

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2412.00061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00061 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:15:07.306486Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:26.218271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:02:34.639740Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7dafc0a5-e07b-4b38-9f4c-963008ed5712 · outbound

This paper cites GPT-4 Technical Report.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.134891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.134891Z digest=sha256:2c90fbc4409d060696ae5452f25a97e3090fc2179c5afc5bf8747e71e7014fed

Observation 81e64fd7-4189-4b95-a1f8-8a6d545713fe · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.141779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.141779Z digest=sha256:0a38bc3b18686295e11d2d3bc337b34cf33e54d4ba62cb8416a760a21cf81e9a

Observation 39c6ce24-f594-4281-973d-e1b8eceb7b43 · outbound

This paper cites Medusa: Simple framework for accelerating llm generation with multiple decoding heads, 2023.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Medusa: Simple framework for accelerating llm generation with multiple decoding heads, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:15:07.802892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:15:07.147755Z digest=sha256:379b963130ef56903fed75d2ac14ce41f65ead76a2476ed438321bf92b47fda7

Observation cdc6218a-69f1-4979-98a0-ea0fcf3165cc · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Accelerating Large Language Model Decoding with Speculative Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.154072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.154072Z digest=sha256:12e4008139c8d3e982e0d8f7e0f9c11b980327d89240f9501cb21328d01c2ac9

Observation dbb37119-5ca2-498f-9683-4fe5acd61444 · outbound

This paper cites Cascade Speculative Drafting for Even Faster LLM Inference.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Cascade Speculative Drafting for Even Faster LLM Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.159423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.159423Z digest=sha256:967da826af3e82e18c2a94bc3e538a6ce3ecc6ffcc31312e854f6577b864e779

Observation ec52fe2a-7e90-4c52-a0e6-0d1e81cbad69 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.165281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.165281Z digest=sha256:df9dc7cd72e4dd902e29c6a9786116253d485b350fded3d8505106f9b2777ada

Observation 6bb13f95-7222-43db-8910-0a07a5aa976d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.170268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.170268Z digest=sha256:1957df96d2b2d7e5173604c2d78a15a860300b8d2d0aaf132a01203e0957b0e3

Observation 680c2458-c6e9-49e8-bac1-419f3190087c · outbound

This paper cites Tutorial on directed acyclic graphs.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Tutorial on directed acyclic graphs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:15:07.772628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:15:07.177796Z digest=sha256:e48f1201dae7913076fdc48b6f4771d23e60bde0e3978621961d4af7d55a12bf

Observation beed94b1-721e-44ad-8974-102fac87eac9 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.183111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.183111Z digest=sha256:f99aab99f3f630d7f78175b76ca56ac3c4252117faa9799216a601c970d0d3d8

Observation 746dfe29-8e6b-4a38-a806-c19cbc1ca527 · outbound

This paper cites SPEED: Speculative Pipelined Execution for Efficient Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration SPEED: Speculative Pipelined Execution for Efficient Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.188194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.188194Z digest=sha256:3b758027d3b41a7d7764206b7b659dca7ba1ecfd9debd21d4c5079ebb43ecf3e

Observation 67459f86-ae2a-4eef-a242-e0fb2c7f3ad4 · outbound

This paper cites Mistral 7B.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.201980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.201980Z digest=sha256:7ab492f4f3138481e5d3646a7e1ef8303e19e8bcd662254ea18af106c5798b4c

Observation 5e41ad1c-e367-4eed-a1f7-21f74b60999f · outbound

This paper cites Fast inference from transformers via speculative decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Fast inference from transformers via speculative decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.207424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.207424Z digest=sha256:e3e9e5ce600e2811783e22fdcb98b68f9ef8fa6647894e1298dbbb7a2ad39b32

Observation 035dd2d0-7333-48d4-8ba3-e175db4f4d79 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.213417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.213417Z digest=sha256:cb923630a099ec238081a7a1dde480d42f1a2d02ec7b3580ad88c481f87df99a

Observation 139f8768-1ac8-486b-a60b-659023c6dd8d · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.218014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.218014Z digest=sha256:7af931415a626d72e30787a09fe495973b8033d731560bb5bfcc60d3e6c5e326

Observation b955a8cc-9cb4-4edf-b883-f1225addf738 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Accelerating LLM Inference with Staged Speculative Decoding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.224514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.224514Z digest=sha256:22d13b27638d382e0c7bc85314c633e2a98e88150c86899d76362447ff889a52

Observation fcf09a45-24f9-42ea-aa36-08e6cec69b61 · outbound

This paper cites Insertion transformer: Flexible sequence generation via insertion operations.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Insertion transformer: Flexible sequence generation via insertion operations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.230200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.230200Z digest=sha256:049bb4afc2a5f3ea27f972857ab5f20b303ed49c6e11aaf872c8c0bd6d075dcf

Observation 8554d824-62b3-4d4b-a860-74ce9a8650a1 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Blockwise parallel decoding for deep autoregressive models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:15:07.713516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:15:07.235539Z digest=sha256:cccf63e681a937d85638aad421ead5c6913eb66ae36f71929422a87a5e7aa4a5

Observation f7555fba-55e2-4eb3-9abc-3605e48fd5ec · outbound

This paper cites Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.240855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.240855Z digest=sha256:9debf4b2e814eb5ce79d8194400c81554cc07317247965e315046c41498dfa0f

Observation 46a3c2dd-d6ee-4b52-b306-a8a8358e1e58 · outbound

This paper cites Spectr: Fast speculative decoding via optimal transport.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Spectr: Fast speculative decoding via optimal transport

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.245355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.245355Z digest=sha256:2d7948074b12654cff01add8fc345c6274288da3646e5fe2454f2ce5f22721f0

Observation 69cd8bde-5543-4faa-8d6b-5828699f8564 · outbound

This paper cites An introduction to conditional random fields.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration An introduction to conditional random fields

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:15:07.682407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:15:07.250401Z digest=sha256:ebe0b5c7dec26c1b40e64d20ec8433a0923e5ce9a565719d85d590b985d7cab4

Observation e5fb8e5f-0bc6-472f-9433-e9c514a86e46 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.256088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.256088Z digest=sha256:5254d56d4d7a25f9e1387f449ef9769a14c2109e6412372ff711e9f3598d1356

Observation 74eb00b7-fb34-424d-898d-79eeedc215a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.262618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.262618Z digest=sha256:052237c4cb5725b3ddfd07684a1a3e6b7ede9d765620d415413fdcf3ed47ccc8

Observation b351093e-64fc-4cbd-a61f-1569b67f3936 · outbound

This paper cites Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:15:07.668268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:15:07.268976Z digest=sha256:40051901958d32e87f2e9595c323b70e007c25bbd8e36659f6b382472624a521

Observation d116a5fe-8de5-42a1-9943-686b508e1d1d · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.274460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.274460Z digest=sha256:30b56b66925523cb84c831098f5e79aa13bbab965bbd160b820dd5f4fbd5c34a

Observation aeb6f735-149c-4e40-895f-ca6eae19d0ba · outbound

This paper cites Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.283186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.283186Z digest=sha256:0700a5de167ad3a6d6cd2efb7e52db792580f7c9c3249af690de218ab58ff79d

Observation e0e5002c-af63-483c-9901-77ee3353e1d0 · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.288948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.288948Z digest=sha256:dd6413398122f05fa0a06a6603f4314ceff0f155a40729cfb758a27a65ea144d

Observation f8a9214b-1cd2-4e06-881f-14e16d536fba · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:15:07.652003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:15:07.301072Z digest=sha256:a0c8f5f77aeb68dff83ffc3ad282e509cd01dd5780cf69f744f815301df8098c

Observation e8bcaf73-917d-478c-8fad-6b701a3a0a92 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.306486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.306486Z digest=sha256:38b1899152d5669103849c2298d436352b2cda3081307f31e460656aeab1e88d

Pith citing papers

Observation dfaa0113-e027-4ed0-9e7f-fb27dead30ad · inbound

Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs cites this paper.

Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:26.218271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:26.218271Z digest=sha256:d8c68533aa87119dde7348a2f46846d3df8008d11f3b115fdbcab6f477b052d3

Observation 9c1fd930-b4d6-41aa-ae7c-cadfb9313ffb · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.641097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:c7ada9b06c66cc6492e850d5da77a396ce2965624300042c790a6a471f52a5ec