Pith. sign in

Paper Citation Record · LEDGER

Faster and Better LLMs via Latency-Aware Test-Time Scaling

As of 13 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.19634.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19634 v4

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:23.666202Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aac5cdee-36ec-4931-83cd-e14e19b84e5d · outbound

This paper cites online" 'onlinestring :=.

Faster and Better LLMs via Latency-Aware Test-Time Scaling online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.849185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.849185Z digest=sha256:bb0b91ec594a122b57bec9556e88ee30e8cc559b18ae6bace23255b3d31078fa

Observation d307be07-e554-4d3e-8550-97389933be11 · outbound

This paper cites write newline.

Faster and Better LLMs via Latency-Aware Test-Time Scaling write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.896928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.896928Z digest=sha256:35b3cb5509209a60fa463d1ce700d7044ef6eb24914f908a05738899798f3c7c

Observation 52ffbfa3-4311-478a-b7b2-70cf8f7e6922 · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.990978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.990978Z digest=sha256:5944088611c45ecd0821647eafb00a61db62f069175aa45bf18c45ad0f208eae

Observation 7c906ed0-9dec-40a0-9909-6be861d3490c · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:25.266685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:21.050136Z digest=sha256:ea5246184880950357f23b85630b23038c05d12a720a04618afa079f78059446

Observation 2475f915-e570-461e-b468-fe9f582b02c5 · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:25.109306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:21.137331Z digest=sha256:e5ce5d32f594e627fa72140ad2aafe156bd14d917b1b96981cfcb748c2b76b33

Observation 83b2dbee-93be-4b24-afe5-63de353efd1c · outbound

This paper cites Qwen Technical Report.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.211170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.211170Z digest=sha256:bc2bd7d069322da9ee6ae8c40c2f7ada532065b1886f2748eab6d5af6d7f8dcd

Observation bd6a1831-87fe-45e4-ad82-3e1bedfc5a75 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.279768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.279768Z digest=sha256:43bc685ef60420ec8833ffccf78e5b4f85e3ae387192657dbfa8cc5e9978c6da

Observation ad607972-c16e-439e-97a6-6881fe0a61a9 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.343342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.343342Z digest=sha256:c3238ca8e2d1400e1ae0240ff8122b42acbb0ab7af539d22b3c44333913011e8

Observation 84de4233-7493-4613-a096-cdbfbc27b29d · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Accelerating Large Language Model Decoding with Speculative Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.428644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.428644Z digest=sha256:28117169c1aef097545c3fe739186cc89112e51a233db0c4f0bfb4340499b6ce

Observation 512295ff-e0ef-409b-94d9-2f00cf25795f · outbound

This paper cites Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.520298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.520298Z digest=sha256:e66adb8c958ec6bc138fb92a6ee01af1340d3dfe781ae79ccc19247e7e0d80db

Observation bf166085-f870-4f4d-bd71-c6c1dacb4a81 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Faster and Better LLMs via Latency-Aware Test-Time Scaling rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.569416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.569416Z digest=sha256:d813b687c86fccfff40a94e99a05f877f6effba535b30367ebdf955f21169968

Observation ba8dbf1b-46b1-445b-9979-ef2715ad2f57 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Faster and Better LLMs via Latency-Aware Test-Time Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.669934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.669934Z digest=sha256:36f3c690a34b29defb2797a75b9e1e2f600e02a79d994cc2c6c24ef8f8028dcb

Observation cbaab37d-1999-4ade-b959-bfd2b75494a2 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.759921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.759921Z digest=sha256:7d81d89cab0c43171837c495fefc565f01ebd668a1e419202fc891fda3debe13

Observation 2a0b5809-9035-4686-85b8-f5284a19deae · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.835787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.835787Z digest=sha256:0804a36ff05f19ca168663bafe26cfcbf6b1926a28ce1a7a527fa43c0acfd0ed

Observation ee65c85e-763f-4d02-8b89-2c59df45ac41 · outbound

This paper cites Evolving Deeper LLM Thinking.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Evolving Deeper LLM Thinking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.933314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.933314Z digest=sha256:51f2ef9217e821103a2b57329a1dfdc50b33673bb9aa5d0a2c102e1c95552f2e

Observation 048e12d2-b409-4c16-833b-57db1ee5873d · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.011774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.011774Z digest=sha256:95eb0ac022cad0227f000e9055cd4af4f203227c80f9070e9f7a6353a2172db7

Observation 62d29a2d-c800-443a-a3d0-89b701e61319 · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:24.941453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:22.100607Z digest=sha256:34d8e4a67d7ab31e53cb7207af7f7ef16a3dec82146de2c5b6b62fdfd5b28ea1

Observation 2edbe66b-5b88-4e7d-ae3b-2b30a219c889 · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:24.773911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:22.192586Z digest=sha256:96dac36420149b5f2d4e3eae88c7ab9af4690d8ba9768f7e9ecb7da07dbba7c3

Observation 006f8452-bf4b-4da0-a950-496d4078bd4a · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

Faster and Better LLMs via Latency-Aware Test-Time Scaling EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.274278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.274278Z digest=sha256:a4304755cf2fd7de5f112ac9303a4fadb4630f325a08b5002ffef8c3fb0c1be7

Observation 770b0951-6d1e-4e04-80c8-1b81e777a340 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.357992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.357992Z digest=sha256:d2406c797035848b62e88628b095a4875d64ede7c8cf0684e5e2c0c5bcb56afd

Observation d68c18c2-584f-499b-8e3d-c4226350db39 · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:24.594446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:22.452686Z digest=sha256:28ffa95ff159a7107fbc12af07340edf5dde39068abfd138076496cc8dfb44ee

Observation d5be18d2-cd9d-4710-909e-f63f96d2644c · outbound

This paper cites s1: Simple test-time scaling.

Faster and Better LLMs via Latency-Aware Test-Time Scaling s1: Simple test-time scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.505414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.505414Z digest=sha256:e3d8ffcf85fcdb08f687e851df742db311d2e2d898ac25e569a9b8822f73682e

Observation 60d564a2-4815-4eae-a3a3-ec135f2e259e · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.568702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.568702Z digest=sha256:e964040abad0b0840f022876f4a95bcdd14b92e1c2d9b258df730bd471ccebfb

Observation d5914a40-4819-44ac-87c7-e1d7cf7beea7 · outbound

This paper cites September 2024.

Faster and Better LLMs via Latency-Aware Test-Time Scaling September 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:24.445470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:22.634383Z digest=sha256:bf951e369c711b39316f6a0c3657c230eb4761289ff2a9c480fe50dbe6561979

Observation 42edd81e-fbf0-4256-8f71-d1834e4d9b6d · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.697519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.697519Z digest=sha256:58d85c0315007ec57cac8baeec438df733364d56de67f847b4baef4725531077

Observation 00262278-648d-401e-83b7-43a01aac0f27 · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.796446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.796446Z digest=sha256:7ccf25d502474e102ab01dcabd323ed233be87f8062d645e094132b1db078ef2

Observation 89eb6f73-dcde-42a1-b3fd-c96b978da84e · outbound

This paper cites Heimdall: test-time scaling on the generative verification.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Heimdall: test-time scaling on the generative verification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.894186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.894186Z digest=sha256:d05b207249323d9fbcf2784f9d533352e0182a5080059f248143c4be86abfdb4

Observation 0d4cf23e-6a88-436b-a0f7-ed0547f2a68d · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.985586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.985586Z digest=sha256:ce91be8176513d3567407760a7742dbd326cb4d4b532353a88d4f497c7e24be3

Observation a7deb134-8cef-40b6-b5d0-244695c3b39d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.077538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.077538Z digest=sha256:7530e2db7a3812efb9666604f89a2d667a18c1f527a377490c56c8767115e517

Observation 41ea4e1f-7967-4149-b3e2-257ab2c6bb58 · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.158186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.158186Z digest=sha256:3a121c2fb3e2d91422e1aa5fd1cb3b1c3e7ab94562b52163f85ba227b273c0b1

Observation 67fdfbae-6c6f-4133-ab19-e02a88109ce1 · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Faster and Better LLMs via Latency-Aware Test-Time Scaling OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.295558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.295558Z digest=sha256:6ba269c3f0f9ad00df30fac168d7285910199bff12b8d2d4204caa648c5f09e9

Observation 4a55ccb2-7b8f-43a4-b605-a5962d289216 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.344701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.344701Z digest=sha256:811680a81b10bc2427d6a7f56fd4e0f771b6adbcab318899397e275728d79bf6

Observation 4b206e90-f942-4f54-b449-7a5212cfd523 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.391333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.391333Z digest=sha256:e38497b8c72b8d33fb4c11ef7d9c1fd014dcc3ba9f49f251e1f8b9e16a4e3299

Observation d5f2f4f7-22bc-4f41-bb88-a4411fed815d · outbound

This paper cites Qwen2.5 Technical Report.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.470946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.470946Z digest=sha256:d312374cb53e830d8a10616ec9d63fa6afe357e7a9f667f1c1fdbb099f3ccc03

Observation c46c84db-aff2-47da-9adb-06ab2e074e8b · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.566465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.566465Z digest=sha256:99235e078cc7e28cf628f181c6a5f29c176b2c535daf0638a64366a8e897f010

Observation edf41162-daa5-4d41-bcd7-c1eeda2c85af · outbound

This paper cites an unresolved cited work.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:24.243516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T14:14:23.610993Z digest=sha256:ed125a63dbca96a4099490ad9485cef593f6cb136b7ea7b5e44406ea90719bb8

Observation b28ec7bd-1095-44a0-8644-3dc96c72738c · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:23.666202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:23.666202Z digest=sha256:8e9b3f0aa9c33a48cfafd53b41be951fe5982a251b7c288a3e02cb53379893a0

Pith citing papers

No inbound Pith citation observations are available.