Pith. sign in

Paper Citation Record · LEDGER

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2509.04474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04474 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:48:23.144764Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0f3bf10-0159-4c65-a2a0-109be6e5d20d · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.159265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.159265Z digest=sha256:1d6efaa3dd33cb3990f2ce1d1361e46a914e426fa1f50d0546e4ba8b28d38db1

Observation fda6280c-cd33-442d-8336-8d4c82be5159 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Training Large Language Models to Reason in a Continuous Latent Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.516547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.516547Z digest=sha256:f105aaf819071db9b6b364fa2cde6056751a3ba6818b5c21f5f5a912b7f24eff

Observation 97d15c73-75eb-468f-b5a4-b949fa073faf · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.574388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.574388Z digest=sha256:6ffb943ad407373a57859461346c8fc29a87af02027a3558df3759f7e7f598a9

Observation e8487933-93ff-4006-9f1e-1d273b3ac010 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Competition-Level Code Generation with AlphaCode

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.656281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.656281Z digest=sha256:e0905e8bee9f196cfb45a9e6e5d8d12141cc711d4afba6b34eb84ca1c1220157

Observation e2392494-a000-412a-a0a1-822dcb38bf9e · outbound

This paper cites Let's Verify Step by Step.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.885995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.885995Z digest=sha256:194398827b21345dd8208aa02723807cf27ea0203e5eb85a1aca9fab7fe6f902

Observation b75ec70c-a409-46f9-af64-e1c149f1ad54 · outbound

This paper cites TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.999552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.999552Z digest=sha256:41d54d60478c63efe94718d2c9b6d36d9e710a42805466607e0084d6b3a3ff4c

Observation 5498e677-3537-4449-a6be-a819a8f98388 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Reasoning Models Can Be Effective Without Thinking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.081867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.081867Z digest=sha256:d59c5fedac5fe5caa0c304621d4f8c336c8cc10ab99890b603e5fa13368c2c8e

Observation 652600af-233b-4e66-b3b0-db8b68ae23de · outbound

This paper cites Suffixdecoding: Extreme speculative decoding for emerging ai applications.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Suffixdecoding: Extreme speculative decoding for emerging ai applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.195104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.195104Z digest=sha256:b4a40c6584ef69ffc4ee87b4fd4e486ae2a4b5f031440ca248f1cf5e05460422

Observation e740dd57-f9fe-439c-b562-be754c3a9e5e · outbound

This paper cites OpenAI o1 System Card.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.337886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.337886Z digest=sha256:96c9b0af5577f62055cccc7868e9944ed3ee9d827c502f0356245a05126fdaa7

Observation 5774eab7-57c0-4983-a394-583f0c9d3f03 · outbound

This paper cites SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.435299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.435299Z digest=sha256:324dadcb031c17e5d629e786938ea07901ec8d12797d3198479df99fc5fed887

Observation d1a4389e-4c0a-4ff9-a3e4-15cf770075bc · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.547651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.547651Z digest=sha256:df099c15afd313dbcd03a415227a999a194677f3872bdd2eb2adab58ad6586a3

Observation fc50e29f-9b0c-4dff-9163-54b06bb2b2c0 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.659104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.659104Z digest=sha256:67faf859821dbc84a6015f3eccb22b1ddb6460aa75b6453a09f014694e0a0263

Observation 248c6c22-6920-42e5-9b6a-c6d3a7e41172 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.756427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.756427Z digest=sha256:9c08fbbae944e573b92a44a36e2d8583fd53734e64fcb639be6228878648a146

Observation 66927755-e89d-481b-97db-e4e91bf8aeeb · outbound

This paper cites R1-compress: Long chain-of-thought compression via chunk compres- sion and search.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling R1-compress: Long chain-of-thought compression via chunk compres- sion and search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.822091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.822091Z digest=sha256:caa7056d861f909f9fd4cf27fa8e4d50cb82a76652e2fd31cf251e0aa60970cf

Observation 178cae99-eb4f-45a1-9c87-f08a431ad679 · outbound

This paper cites Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.954670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.954670Z digest=sha256:bcad4312785313249bc82486c3165e64707f1ebc3427cdf2f35d719c3b266e5d

Observation d24ca8c5-ad1c-4519-b0ba-d3b3ceea96cf · outbound

This paper cites Qwen3 Technical Report.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:23.073663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:23.073663Z digest=sha256:ee554696c52c6e3b4b97612fa8df80c2040ad0bb00d1407c8210bd87ae803d61

Observation 5ff7f1ea-4082-41c1-90fb-2770e16904bf · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:23.144764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:23.144764Z digest=sha256:37550d95000ccb1c34538ac4ede361f44a419931b24b5292dffeef7afe40f38c

Observation 393a2bd1-c2b2-4eab-a223-8903dcaf506e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.341874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.341874Z digest=sha256:aa816369d842e623eb798506404e134936a1477a20e07e8f8adf29f9c4106cee

Observation 591a950f-0fcf-4023-b754-e1aa1b44acc3 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.764917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.764917Z digest=sha256:6854cae30a4d20fcbb1f24ebcb6b5c3b11f44afb609ef31befaddc1e082155d3

Observation eb14ed75-474b-4cd7-87a3-cce4ce4bebb3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.248460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.248460Z digest=sha256:2cdcdf666939106b461fbf49b6983a369ca6bd0cf79a2b5eee91e1da7173444f

Observation ef734227-f47c-4a37-9275-0c36f298f419 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.101641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.101641Z digest=sha256:7094a2e2427d0babcd492fbcc68f8315f59ae4ecfae2f9b94d3ef21271e1ed8b

Observation c386ed0c-2dca-44de-98e1-0ab2d484f2ae · outbound

This paper cites Break the Sequential Dependency of LLM Inference Using Lookahead Decoding.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.420932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.420932Z digest=sha256:97835089d0375dc882f9a90f20264a5239f36ada4084575d0eb9e07dc7302e68

Pith citing papers

No inbound Pith citation observations are available.