Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:38:54.851579Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 4 inbound Pith citation observations for arXiv:2508.08192.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:38:54.851579Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:15:30.174288Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T22:54:09.662817Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0af16ab2-4b16-4e29-8b75-adef00a83fb5 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa6a4eb-df24-455e-b6e1-dac3ab84623a · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The curious case of neural text degeneration
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4bd292e-1473-4a07-954c-7d0590f9378e · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Hydragen: High-Throughput LLM Inference with Shared Prefixes
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c659adac-c136-43a8-a190-429d4a696843 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Adam: A Method for Stochastic Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22dc4cd0-3f2c-4a13-9d49-ec375612ab34 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def561ff-a131-44a3-93d8-5e6c10b04c7e · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235d3e2c-a8ac-4b40-b616-9ae04db67d3b · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4fb8f7-47d7-454d-97a3-2158615ea817 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e518d1-e991-443a-a8e0-5e07929bd86f · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Self-attention Does Not Need $O(n^2)$ Memory
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e294fe4-01d2-4df9-951e-e13471fef135 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a069bcfd-b7fd-43b0-8302-5ec377fbd2f5 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5258b7-e136-43be-956e-618bcf07e180 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Efficient Guided Generation for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13dfcad-6739-4a19-afb7-61714e8215ec · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The Synergy of Speculative Decoding and Batching in Serving Large Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907cf67f-c951-468a-9f60-d087a4d7a53d · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 828ff934-d954-451a-aeba-d4f35015f049 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d9578e-0d0c-490a-b387-ee34a99bc436 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Flash-decoding for long-context inference.https: //crfm.stanford.edu/2023/10/12/flashdecoding.html, October 12
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a254f1c7-000d-4e24-aff4-5321cd42a9bb · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions The Llama 3 Herd of Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ef08b4-b26e-4bd0-875b-27f2784c268e · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Accelerating Large Language Model Decoding with Speculative Sampling
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee79eeae-b9dc-4d04-863c-a77925ad01d5 · outbound
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Xupeng Miao et al
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17d775b5-4e95-4970-b31b-6bcf50c3f27d · inbound
HiSpec: Hierarchical Speculative Decoding for LLMs Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e20ca68-c9a4-48a1-b6e0-3bd8ab542d69 · inbound
Speculative Decoding with a Speculative Vocabulary Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8e80813-20fa-4aa2-8701-38fcc58e3bda · inbound
Test-Time Speculation Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e90e11a9-e217-4497-9262-f8b754177ac4 · inbound
Test-Time Speculation Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.