Pith. sign in

Paper Citation Record · LEDGER

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

As of 13 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 4 inbound Pith citation observations for arXiv:2508.08192.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08192 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:38:54.851579Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:15:30.174288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T22:54:09.662817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0af16ab2-4b16-4e29-8b75-adef00a83fb5 · outbound

This paper cites GPT-4 Technical Report.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.423030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.423030Z digest=sha256:40f450d7535eefd2c5e94965cc2f0922294388101e952fea42bc1182c0954613

Observation 2fa6a4eb-df24-455e-b6e1-dac3ab84623a · outbound

This paper cites The curious case of neural text degeneration.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The curious case of neural text degeneration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:38:55.506409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:38:52.897703Z digest=sha256:6756341849137ba2c506fabb9a59a1a1a6110522c59626a12390f07bd8fa550e

Observation d4bd292e-1473-4a07-954c-7d0590f9378e · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.050106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.050106Z digest=sha256:de3f781519a2aef58ee238a14dfb0a13084dda3f88617fee786373597c1c59af

Observation c659adac-c136-43a8-a190-429d4a696843 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Adam: A Method for Stochastic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.157905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.157905Z digest=sha256:2212bdda238acb9933e21fe9282e6c75c158af5d9eeb046ff3cd287b8eb640f0

Observation 22dc4cd0-3f2c-4a13-9d49-ec375612ab34 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.285462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.285462Z digest=sha256:05c4f2764d7ac3339d951abcbaa817dbe6d18bbdd551aef69a4af9b2169d5f3d

Observation def561ff-a131-44a3-93d8-5e6c10b04c7e · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.476154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.476154Z digest=sha256:23d06dcbf11c32599d9585d07b843d51b53dc84fd5e3fc912aad38bf46597a99

Observation 235d3e2c-a8ac-4b40-b616-9ae04db67d3b · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.660836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.660836Z digest=sha256:625e00166bc3bbafefdbe920d67fb6e556ed1b30d68f5ef3827445375d20cb03

Observation 7d4fb8f7-47d7-454d-97a3-2158615ea817 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.075982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.075982Z digest=sha256:ea2a65f246f5aa6f08b6fc8f335a9298f0aa468e22634641cab14f4c6ced68e2

Observation 99e518d1-e991-443a-a8e0-5e07929bd86f · outbound

This paper cites Self-attention Does Not Need $O(n^2)$ Memory.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Self-attention Does Not Need $O(n^2)$ Memory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.293572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.293572Z digest=sha256:37ea32e6303e877fdc4a911b41b03d592883aaee29cbc32951ef78621073622b

Observation 4e294fe4-01d2-4df9-951e-e13471fef135 · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.422691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.422691Z digest=sha256:79242b5ef32f1fc54ca6b2713b21ed019f5436fb4a8ee12ca4777924af5592fd

Observation a069bcfd-b7fd-43b0-8302-5ec377fbd2f5 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.550812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.550812Z digest=sha256:e7ef44a7601c347f2de262a3b3e3f8c52e73820a99d3edae801d0e95f88a7267

Observation bb5258b7-e136-43be-956e-618bcf07e180 · outbound

This paper cites Efficient Guided Generation for Large Language Models.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Efficient Guided Generation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.851579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.851579Z digest=sha256:e55332471c897701089b6b381f58a649b21a6520a1d145a052661415117207a1

Observation d13dfcad-6739-4a19-afb7-61714e8215ec · outbound

This paper cites The Synergy of Speculative Decoding and Batching in Serving Large Language Models.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The Synergy of Speculative Decoding and Batching in Serving Large Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.729462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.729462Z digest=sha256:5beda2cdaa57c7c650d2c7249cb5125d9f2e1970e04d3756f246787a2b219863

Observation 907cf67f-c951-468a-9f60-d087a4d7a53d · outbound

This paper cites Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.973941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.973941Z digest=sha256:33a3691aa37ce771829b40db003a51c0d26359b507d1bcac094cedc0576e1307

Observation 828ff934-d954-451a-aeba-d4f35015f049 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.474661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.474661Z digest=sha256:398d6931ea57a14e551ba7d6e7c246bf70280c95a6ae171987d3fb355ca28ce5

Observation a4d9578e-0d0c-490a-b387-ee34a99bc436 · outbound

This paper cites Flash-decoding for long-context inference.https: //crfm.stanford.edu/2023/10/12/flashdecoding.html, October 12.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Flash-decoding for long-context inference.https: //crfm.stanford.edu/2023/10/12/flashdecoding.html, October 12

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:38:55.702221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:38:52.670091Z digest=sha256:45112221e1a85deee400da34c94a48b931d7f96a4444d8ac034a3ece55294174

Observation a254f1c7-000d-4e24-aff4-5321cd42a9bb · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.767983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.767983Z digest=sha256:3057e8e4992c079ae0644b808d41b8fcd0b519d8b866ab147ebee5fe79a131cb

Observation 33ef08b4-b26e-4bd0-875b-27f2784c268e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.571364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.571364Z digest=sha256:8dba5ddf39f52e002ba645e73ce13bf7c0d6e18b42a9cacf7a470a1d2052fabf

Observation ee79eeae-b9dc-4d04-863c-a77925ad01d5 · outbound

This paper cites Xupeng Miao et al.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Xupeng Miao et al

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:38:55.238570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:38:53.815994Z digest=sha256:9264c3b49cd01020f4e88a6823a4a693f674eb881953970847b7ed90fb510991

Pith citing papers

Observation 17d775b5-4e95-4970-b31b-6bcf50c3f27d · inbound

HiSpec: Hierarchical Speculative Decoding for LLMs cites this paper.

HiSpec: Hierarchical Speculative Decoding for LLMs Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:30.174288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:30.174288Z digest=sha256:84988578948809727091c085a9abdac2cfaf2278913aac953977db94ce3a19cf

Observation 4e20ca68-c9a4-48a1-b6e0-3bd8ab542d69 · inbound

Speculative Decoding with a Speculative Vocabulary cites this paper.

Speculative Decoding with a Speculative Vocabulary Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T23:27:56.589526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:27:56.589526Z digest=sha256:cb09309a2f7a279e7891cd7b3fa8131d64c49a32cafa965fdadd907b2501e77c

Observation e8e80813-20fa-4aa2-8701-38fcc58e3bda · inbound

Test-Time Speculation cites this paper.

Test-Time Speculation Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.742941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T04:20:15.459967Z digest=sha256:f18c4a967b2d31d0f03f95a651ec8b76c059ca5056ed4e51df01e33ee4ec145d

Observation e90e11a9-e217-4497-9262-f8b754177ac4 · inbound

Test-Time Speculation cites this paper.

Test-Time Speculation Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:54:09.665858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-20T22:53:56.914556Z digest=sha256:a48d82d4fc2da527abe4657a44b635e0ed127254455911b35ca6825fe3df6261