Pith. sign in

Paper Citation Record · LEDGER

Test-Time Scaling with Reflective Generative Model

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.01951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01951 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:46:44.808617Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 649a06c3-d709-49f4-8d49-07364382960c · outbound

This paper cites Accessed: 2024-12-20.

Test-Time Scaling with Reflective Generative Model Accessed: 2024-12-20

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:46:45.017812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:46:44.756380Z digest=sha256:0240a93951f7a2e6e5565a5144d052b54af40eb263fa45c3f87e7fc13e315c77

Observation 8226f95e-f046-4cbf-919f-72c19f342603 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Test-Time Scaling with Reflective Generative Model rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.759235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.759235Z digest=sha256:00174da85958d92131b1dd7d9cf76dfed73cdf9b1a758cad34f0a3507db31813

Observation ffd2d00c-1b1d-470f-a630-a5176093bc0e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Test-Time Scaling with Reflective Generative Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.762329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.762329Z digest=sha256:96c387f817b006c3bb2be7946a40f3bee4cdfeebe623aba5930517bb264593d2

Observation 8385cb18-d209-40df-84d9-676f3c50eba2 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Test-Time Scaling with Reflective Generative Model LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.767588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.767588Z digest=sha256:76a13f026f73ef0558d47afc48325a74231fa37380d0284ace13fe09a8230e05

Observation 6cb25bc9-63f5-4393-a4bc-dee21072b0d2 · outbound

This paper cites The Impact of Reasoning Step Length on Large Language Models.

Test-Time Scaling with Reflective Generative Model The Impact of Reasoning Step Length on Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.770899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.770899Z digest=sha256:ba7f4cc38d1b94c5c37ec35ecbf225087cf4dfad68800d601ae8db6599469489

Observation 351255ef-6d68-4d7a-9219-6e99ab6ba5a9 · outbound

This paper cites Accessed: 2025-02-18.

Test-Time Scaling with Reflective Generative Model Accessed: 2025-02-18

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:46:45.009803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:46:44.773451Z digest=sha256:0eae7f608788c3b6bc80c625b5a773559d5f9a04902c1583105aecf546a7ff15

Observation 960841b0-a8eb-43b4-ae90-d6e7ef6b0273 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Test-Time Scaling with Reflective Generative Model LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.775814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.775814Z digest=sha256:95728bc72487ef4d399e40768e40deceaa95532a31bc65f84f7d99d433aec0ec

Observation 47967841-d013-48fe-a5c4-a9853eeb27be · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Test-Time Scaling with Reflective Generative Model Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.778485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.778485Z digest=sha256:8256407b2cf2df524122b2b650fcc3b15371a78d4f1a736763fad188d30fe326

Observation 7ebdc463-bd03-4a77-8dc5-8396593fd947 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Test-Time Scaling with Reflective Generative Model Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.780934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.780934Z digest=sha256:50b17739255c29eb61d4042cc51e750af7f7803f140833fef170e1d456bb652a

Observation 1845ae5b-eaf4-43c4-a421-8b0b14f25fd2 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Test-Time Scaling with Reflective Generative Model Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.783885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.783885Z digest=sha256:2f06d28aa722cff0c6deb86d41b71d5c4b10f343bae2795db928355ea474d3d9

Observation d3880fce-8e76-4f62-a8e7-1d6e0efecb7b · outbound

This paper cites s1: Simple test-time scaling.

Test-Time Scaling with Reflective Generative Model s1: Simple test-time scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.786503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.786503Z digest=sha256:bf41d49dd70dab0c31eb45605c9905db8f31023cbee1fe3b7705e3c32c8a2ec2

Observation 2e4d49a9-e75a-4825-9437-4a33a5296366 · outbound

This paper cites Accessed: 2025-04-13.

Test-Time Scaling with Reflective Generative Model Accessed: 2025-04-13

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:46:45.001249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:46:44.789082Z digest=sha256:16b9545521604ce44d664564fca024c4025172f8d8307b5acb8d978655aaf4d1

Observation a4aead36-ea7f-4f24-8df1-5142667e29da · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Test-Time Scaling with Reflective Generative Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.791690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.791690Z digest=sha256:d3b367b8d0b16f73df306c69a3bd410e482173ad954117089c9ebec1fb07b09f

Observation 92296b9a-402d-4b5e-8175-2c9887416d9d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Scaling with Reflective Generative Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.794776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.794776Z digest=sha256:f3ccc438d4ae02b956303b993f36ea3fe449905967eece33dc2d92a2e8a48e8b

Observation d3461d8a-0174-4f13-a6fc-b94a0a92ee26 · outbound

This paper cites AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification.

Test-Time Scaling with Reflective Generative Model AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.797312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.797312Z digest=sha256:613cdb87b051f37c440d8d6420fcde8f967b60ac3df223bc3293de4317b1e35a

Observation f5d6b8b6-6b14-4d25-831d-644c5316d7c7 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Test-Time Scaling with Reflective Generative Model Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.799931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.799931Z digest=sha256:328f23cd7aa14fe1d3b1b0a1aef135ae26a55fe332cfd83dd6b23551a03d1183

Observation 85042b7b-9067-4a7b-8151-9a268dc0ed74 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Test-Time Scaling with Reflective Generative Model Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.802656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.802656Z digest=sha256:14017aeef53ff0d9ea32844a9b5ed585c891088c849f304af28fd057639429fb

Observation 54db9098-0616-4541-9ef3-e296ed949322 · outbound

This paper cites Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?.

Test-Time Scaling with Reflective Generative Model Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.805597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.805597Z digest=sha256:7a9ec310cb8c30924ceec4ddc989bd83e3adf20cb11c17ce5cb76ad0f5f61593

Observation b537c84a-e98d-491f-97a5-ad21595fab3e · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Test-Time Scaling with Reflective Generative Model The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.808617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.808617Z digest=sha256:20009fb3fd2666a954386e75dc323bcc4c8707a7b9d3e3d015eb7999351bffeb

Observation efa51012-8f6d-4d4e-9668-632be0b1bbdb · outbound

This paper cites OpenAI o1 System Card.

Test-Time Scaling with Reflective Generative Model OpenAI o1 System Card

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.765061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.765061Z digest=sha256:e635fee35f84e32dffa7042cdf3d4e330844187b669736d7f7449dc15ffe46b6

Observation e9214b3c-3bd2-4c2a-b98c-8965c9052472 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Test-Time Scaling with Reflective Generative Model Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.752442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.752442Z digest=sha256:f7d1c2d54a0144153e116e24b618b5ffe976422dd7746d5d723cd70a3390238c

Observation b532d1d6-c398-4e00-81bd-ba0d9d30f022 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Test-Time Scaling with Reflective Generative Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.748925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.748925Z digest=sha256:14b964483d82f48240aad61cbbcb7ad41a7ded135f1fb2a47a2aecc1d8c503a5

Pith citing papers

No inbound Pith citation observations are available.