Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2405.03553.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:42:44.023438Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:59:38.178795Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 95df9dbc-b51e-4716-84d4-a1ff874671d7 · inbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e7de639-b7b9-4499-b210-0a96af1fd149 · inbound
The Lessons of Developing Process Reward Models in Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aba8b193-025b-4266-b179-838f641dfc04 · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models AlphaMath Almost Zero: Process Supervision without Process
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d942bd42-d465-408f-8559-ec408d011617 · inbound
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training AlphaMath Almost Zero: Process Supervision without Process
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation af617fc0-5954-4655-8222-a5b29688ce9e · inbound
Reward-Guided Speculative Decoding for Efficient LLM Reasoning AlphaMath Almost Zero: Process Supervision without Process
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ff0d55-9264-4da6-917e-1d6b3c42b054 · inbound
Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aaf91ed-6d16-4599-89f9-4da1e791b91a · inbound
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information AlphaMath Almost Zero: Process Supervision without Process
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f524a81-af53-4db3-a350-7734de9c90ef · inbound
Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking AlphaMath Almost Zero: Process Supervision without Process
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cea037dc-7b08-46f4-8717-cedc7b7d32eb · inbound
Iterative Deepening Sampling as Efficient Test-Time Scaling AlphaMath Almost Zero: Process Supervision without Process
Reference 463
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db981569-bbae-4bbb-9a6c-4b523ed8decb · inbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2402dfe1-17ca-45d9-8aa6-e55610864943 · inbound
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition AlphaMath Almost Zero: Process Supervision without Process
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72aa303b-4f83-4ebd-941d-978a51fbb63b · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models AlphaMath Almost Zero: Process Supervision without Process
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6f6a7e9-f55a-4881-9803-c1929b3d812e · inbound
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5deda4ba-41c6-44cd-8c92-812854e08f0d · inbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning AlphaMath Almost Zero: Process Supervision without Process
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4114fc-0234-43e5-b9a5-0942a0a263ee · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey AlphaMath Almost Zero: Process Supervision without Process
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 562b797c-8f13-438d-bcde-0e2872195c49 · inbound
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents AlphaMath Almost Zero: Process Supervision without Process
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1cd788-394a-43d5-bbad-e5da891f9fc6 · inbound
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents AlphaMath Almost Zero: Process Supervision without Process
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb1b3b2-3966-48c9-8a02-100d9ca0a134 · inbound
CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37859dfb-bf4d-40f2-b6a5-eede338bbcf6 · inbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search AlphaMath Almost Zero: Process Supervision without Process
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263e6914-3671-4b38-9243-004647ca830f · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence AlphaMath Almost Zero: Process Supervision without Process
Reference 233
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79e7014c-c445-4614-b3f4-abc9e3f31819 · inbound
Teaching Language Models To Gather Information Proactively AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3751ea62-5e02-48cc-bd39-ce387294c4a2 · inbound
Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling AlphaMath Almost Zero: Process Supervision without Process
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e4ebbc-7384-42e9-ad58-1a12d88c7053 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey AlphaMath Almost Zero: Process Supervision without Process
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da57585d-8e1f-4826-b173-bb9b6272ace1 · inbound
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework AlphaMath Almost Zero: Process Supervision without Process
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ecd80a-1fbe-4573-bfbf-687568a7f43c · inbound
Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space AlphaMath Almost Zero: Process Supervision without Process
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b7865a-05ed-4ac7-96ad-f9356ecc90d4 · inbound
Efficient Process Reward Modeling via Contrastive Mutual Information AlphaMath Almost Zero: Process Supervision without Process
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e1732b0c-e87f-47b7-95d4-6a4c3984777f · inbound
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces AlphaMath Almost Zero: Process Supervision without Process
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0f900f5-924b-43e0-9611-b935c93847ef · inbound
VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models AlphaMath Almost Zero: Process Supervision without Process
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc201724-9778-4f9a-ba21-49f95a24563e · inbound
Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents AlphaMath Almost Zero: Process Supervision without Process
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.