Pith. sign in

Paper Citation Record · LEDGER

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

As of 9 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 4 inbound Pith citation observations for arXiv:2508.08192.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08192 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:38:54.851579Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:15:30.174288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T22:54:09.662817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0af16ab2-4b16-4e29-8b75-adef00a83fb5 · outbound

This paper cites GPT-4 Technical Report.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.423030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.423030Z digest=sha256:efcb4d5e194c73a50f6e6b58bdb75a594eda27574b06463ccd204cda3c817a9b

Observation 2fa6a4eb-df24-455e-b6e1-dac3ab84623a · outbound

This paper cites The curious case of neural text degeneration.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The curious case of neural text degeneration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:38:55.506409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T21:38:52.897703Z digest=sha256:1eb942710e3625ae52e395b15f5b585a16ad6a2f8fade63ad600e365ae0059fe

Observation d4bd292e-1473-4a07-954c-7d0590f9378e · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.050106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.050106Z digest=sha256:3c75a9002e7519a087dd85054ad3389d2bbfcb0f37736e5658ff6294c2b2ba3b

Observation c659adac-c136-43a8-a190-429d4a696843 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Adam: A Method for Stochastic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.157905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.157905Z digest=sha256:175d07e201513b176eec658ba589eae557866a09609fbb54255a266126e416c0

Observation 22dc4cd0-3f2c-4a13-9d49-ec375612ab34 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.285462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.285462Z digest=sha256:5d63e332f250ad56f89256294b8369ff3c28394d9486b011c5124bdcc76f0ed0

Observation def561ff-a131-44a3-93d8-5e6c10b04c7e · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.476154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.476154Z digest=sha256:66bee443738358def9fd144aba2b92a10d943898b33874ebc92df2880add579b

Observation 235d3e2c-a8ac-4b40-b616-9ae04db67d3b · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:53.660836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:53.660836Z digest=sha256:802cc1de700c6f5226237216fcebf35a7fcc4b9c169d06a8e7e28916db13eeaf

Observation 7d4fb8f7-47d7-454d-97a3-2158615ea817 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.075982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.075982Z digest=sha256:08502ee48039de0908b95dfcbf7be21b7988e6704b38de2c8aad5cc4429979b3

Observation 99e518d1-e991-443a-a8e0-5e07929bd86f · outbound

This paper cites Self-attention Does Not Need $O(n^2)$ Memory.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Self-attention Does Not Need $O(n^2)$ Memory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.293572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.293572Z digest=sha256:2f7d73618542f6d16fb1b871ae40ab685f21bd6473d96c5c73f544a8867ccfa5

Observation 4e294fe4-01d2-4df9-951e-e13471fef135 · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.422691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.422691Z digest=sha256:9daf8ef17a6d1c33146345b6f1f05bb4db7b526e587e2dc88ccee112c0372c50

Observation a069bcfd-b7fd-43b0-8302-5ec377fbd2f5 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.550812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.550812Z digest=sha256:946063c14d2af44225b144bee1230547cc1448e1ef83c6a12e533bb4c6aaaed5

Observation bb5258b7-e136-43be-956e-618bcf07e180 · outbound

This paper cites Efficient Guided Generation for Large Language Models.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Efficient Guided Generation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.851579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.851579Z digest=sha256:15d7fb90540da714617975758689483429fc6350e773a98a40cf3be08eb3aa35

Observation d13dfcad-6739-4a19-afb7-61714e8215ec · outbound

This paper cites The Synergy of Speculative Decoding and Batching in Serving Large Language Models.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The Synergy of Speculative Decoding and Batching in Serving Large Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.729462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.729462Z digest=sha256:cbbfc1d8f57995f1dccaeb4e4b782afb4ae21fa88688a330c452e846c1c4bcda

Observation 907cf67f-c951-468a-9f60-d087a4d7a53d · outbound

This paper cites Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.973941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.973941Z digest=sha256:9ec615fb1df14d244091de8f049a3df4c5bb5a206ff77ef908e172ade13ffede

Observation 828ff934-d954-451a-aeba-d4f35015f049 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.474661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.474661Z digest=sha256:4ddc876e23b5bf2a4948230e49ed1025dac0ddaa60d34091ef7738d69229a321

Observation a4d9578e-0d0c-490a-b387-ee34a99bc436 · outbound

This paper cites Flash-decoding for long-context inference.https: //crfm.stanford.edu/2023/10/12/flashdecoding.html, October 12.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Flash-decoding for long-context inference.https: //crfm.stanford.edu/2023/10/12/flashdecoding.html, October 12

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:38:55.702221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T21:38:52.670091Z digest=sha256:2a4bce490ba48966cafbeebc8f3293001b7e902a0c07e0f69af3e112f1c5b8f0

Observation a254f1c7-000d-4e24-aff4-5321cd42a9bb · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.767983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.767983Z digest=sha256:40530879c78e88810d766a587de4f9f0e7bf4143762af55ce45f5c8ba0871d69

Observation 33ef08b4-b26e-4bd0-875b-27f2784c268e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:52.571364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:52.571364Z digest=sha256:c962f3ada5d196b4703afe534a78fc4274d7636f737cef400f265c4fe4a33e30

Observation ee79eeae-b9dc-4d04-863c-a77925ad01d5 · outbound

This paper cites Xupeng Miao et al.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Xupeng Miao et al

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:38:55.238570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T21:38:53.815994Z digest=sha256:0b07fdd49cdba19c2c2ebd258c466bbc526eb0071a065a32fe467421c8f34515

Pith citing papers

Observation 17d775b5-4e95-4970-b31b-6bcf50c3f27d · inbound

HiSpec: Hierarchical Speculative Decoding for LLMs cites this paper.

HiSpec: Hierarchical Speculative Decoding for LLMs Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:30.174288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:30.174288Z digest=sha256:a308b9aeeb011a223cb55dfae736c1bd9f447e021ed54d0be32191bc4cc5a45b

Observation 4e20ca68-c9a4-48a1-b6e0-3bd8ab542d69 · inbound

Speculative Decoding with a Speculative Vocabulary cites this paper.

Speculative Decoding with a Speculative Vocabulary Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T23:27:56.589526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:27:56.589526Z digest=sha256:25b48946032941ce98fad5a5feab1d3c76c39d496508634ddedc98c75aef2a2f

Observation e8e80813-20fa-4aa2-8701-38fcc58e3bda · inbound

Test-Time Speculation cites this paper.

Test-Time Speculation Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.742941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:20:15.459967Z digest=sha256:f702a76b6986590646782055357f1329a90f50712482eb2532f0f74b100cbe5b

Observation e90e11a9-e217-4497-9262-f8b754177ac4 · inbound

Test-Time Speculation cites this paper.

Test-Time Speculation Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:54:09.665858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T22:53:56.914556Z digest=sha256:72486837fae676208d8c84b3e770dda4253916f08aaed07b9f6d93b7e7f17745