Pith. sign in

Paper Citation Record · LEDGER

AlphaMath Almost Zero: Process Supervision without Process

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2405.03553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.03553 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:42:44.023438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:59:38.178795Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 95df9dbc-b51e-4716-84d4-a1ff874671d7 · inbound

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs cites this paper.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.092722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:2bb84a08680b997e163e27e1644552f5561271babc3371e60741e0a049222751

Observation 4e7de639-b7b9-4499-b210-0a96af1fd149 · inbound

The Lessons of Developing Process Reward Models in Mathematical Reasoning cites this paper.

The Lessons of Developing Process Reward Models in Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:43:43.262923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T13:43:43.223872Z digest=sha256:77468068cf9c7f216bd811e4930c0d0a5f7ddea64c1a041e7d82c9b6ced08517

Observation aba8b193-025b-4266-b179-838f641dfc04 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models AlphaMath Almost Zero: Process Supervision without Process

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.482839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:a5c14ea9b853af20e7681262db995d79a47002a4b4bd509425c21001149d5ed2

Observation d942bd42-d465-408f-8559-ec408d011617 · inbound

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training cites this paper.

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training AlphaMath Almost Zero: Process Supervision without Process

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:31:30.251450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T21:31:30.202477Z digest=sha256:b1a97bd9a01b6365fc730072794dba2cd12415aee7b9b164c28fcb89c53aa15f

Observation af617fc0-5954-4655-8222-a5b29688ce9e · inbound

Reward-Guided Speculative Decoding for Efficient LLM Reasoning cites this paper.

Reward-Guided Speculative Decoding for Efficient LLM Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T20:42:44.023438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:42:44.023438Z digest=sha256:301cd3bec415a3b856d93e6e4a042766422280d5125112ea54ec10b5190dd47f

Observation 85ff0d55-9264-4da6-917e-1d6b3c42b054 · inbound

Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning cites this paper.

Scaling Flaws of Verifier-Guided Search in Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:40:23.078834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:40:23.078834Z digest=sha256:6f9889630a7bc122df225602577882d0bf5eb49a512274b2ab78c4bddb3749bf

Observation 3aaf91ed-6d16-4599-89f9-4da1e791b91a · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information AlphaMath Almost Zero: Process Supervision without Process

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:51.975606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:51.975606Z digest=sha256:25fee00adffba964641e7f64ceac390e3732c3a640aa4f66c5b1373233118277

Observation 7f524a81-af53-4db3-a350-7734de9c90ef · inbound

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking cites this paper.

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking AlphaMath Almost Zero: Process Supervision without Process

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T21:40:26.778702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:40:26.778702Z digest=sha256:5859ca0b629177f10ad25998d04e2ec20b14d5ffd03803bbe2099ad6803e907b

Observation cea037dc-7b08-46f4-8717-cedc7b7d32eb · inbound

Iterative Deepening Sampling as Efficient Test-Time Scaling cites this paper.

Iterative Deepening Sampling as Efficient Test-Time Scaling AlphaMath Almost Zero: Process Supervision without Process

Reference 463

Resolution
unresolved
no resolver link, observed 2026-08-08T19:23:26.279454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:23:26.279454Z digest=sha256:cdd95bfe0dc84670955bcbee8e9ef86198670db3d6eae3026b3ced1da2aa2480

Observation db981569-bbae-4bbb-9a6c-4b523ed8decb · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.351858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.351858Z digest=sha256:62994adb8b8d7ae235b73b5da369662a8132f8ce994ee8c3d6098f1505fbc991

Observation 2402dfe1-17ca-45d9-8aa6-e55610864943 · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition AlphaMath Almost Zero: Process Supervision without Process

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.308231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.308231Z digest=sha256:a1a091599a279ec2debd9b0e9132340720a23bf9b39776ec125ff0cea4a1512d

Observation 72aa303b-4f83-4ebd-941d-978a51fbb63b · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models AlphaMath Almost Zero: Process Supervision without Process

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.317561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:19d23a5ee5da923a9ead6ced42358f5daddfc897f38aa2be67f9aa6e9815782d

Observation e6f6a7e9-f55a-4881-9803-c1929b3d812e · inbound

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning cites this paper.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.934365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.934365Z digest=sha256:85f791e206bcbb2550b145256b6360ce9f383126f81fce83c6bf5f05756e4590

Observation 5deda4ba-41c6-44cd-8c92-812854e08f0d · inbound

Reward Model Generalization for Compute-Aware Test-Time Reasoning cites this paper.

Reward Model Generalization for Compute-Aware Test-Time Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.432302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.432302Z digest=sha256:40d6efbf8f638f44e7fd0cb72a60f743444241dd5b6653522eeb79ca75eb7049

Observation 9b4114fc-0234-43e5-b9a5-0942a0a263ee · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey AlphaMath Almost Zero: Process Supervision without Process

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:52.046023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:52.046023Z digest=sha256:304b25988b16f3bde4fca14c915aed28753e0da03548b817bcd5e3253e383559

Observation 562b797c-8f13-438d-bcde-0e2872195c49 · inbound

Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents cites this paper.

Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents AlphaMath Almost Zero: Process Supervision without Process

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.356293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:08:09.356293Z digest=sha256:2f52ba4e7bc2a30c5b7de4bc71604b6f9a684d07083444ea9e943ffb1074ca1b

Observation 0d1cd788-394a-43d5-bbad-e5da891f9fc6 · inbound

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents cites this paper.

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents AlphaMath Almost Zero: Process Supervision without Process

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:12.576689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:34:12.576689Z digest=sha256:342b5ac18c5b98c08b19d5c02530f7439b8c16060ac6a96189007476801ed41c

Observation 7fb1b3b2-3966-48c9-8a02-100d9ca0a134 · inbound

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning cites this paper.

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:38.439633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:38.439633Z digest=sha256:05992838bc8a4945835abd2bbaba1e8adee12df34e4e8bab7eae01701a9e005e

Observation 37859dfb-bf4d-40f2-b6a5-eede338bbcf6 · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search AlphaMath Almost Zero: Process Supervision without Process

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:20.439050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:20.439050Z digest=sha256:e8620a13c163979545e398078daf7b78e1405800d6e8334f1b8df622a708908f

Observation 263e6914-3671-4b38-9243-004647ca830f · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence AlphaMath Almost Zero: Process Supervision without Process

Reference 233

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.196329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:01358b8c408a6690dcc5156624a2ff9dabc16305b5a978b5baedfe016dc579a9

Observation 79e7014c-c445-4614-b3f4-abc9e3f31819 · inbound

Teaching Language Models To Gather Information Proactively cites this paper.

Teaching Language Models To Gather Information Proactively AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:53:23.931239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:53:23.931239Z digest=sha256:7f6097b8c9bf06c68b39751aa475d51e0c2e40b2290c13f0f716d8e2c912c0d9

Observation 3751ea62-5e02-48cc-bd39-ce387294c4a2 · inbound

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling cites this paper.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling AlphaMath Almost Zero: Process Supervision without Process

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.465741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.465741Z digest=sha256:8d029cbb5bc17134cc6dc268c0be0b4f0011d34f5f25dbce21d3f198a46f99cd

Observation 65e4ebbc-7384-42e9-ad58-1a12d88c7053 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey AlphaMath Almost Zero: Process Supervision without Process

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.105505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:ddc1d0df7b8686f59640f422b2e7f4ba07e07eeae34bbd3d2e88af1e3005036b

Observation da57585d-8e1f-4826-b173-bb9b6272ace1 · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework AlphaMath Almost Zero: Process Supervision without Process

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:02.225374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:02.225374Z digest=sha256:e4ea02049648b57fd86a1d391a3553b837c4c31396114c3fd25c6a721444350a

Observation 70ecd80a-1fbe-4573-bfbf-687568a7f43c · inbound

Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space cites this paper.

Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space AlphaMath Almost Zero: Process Supervision without Process

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T07:08:32.399343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:08:32.399343Z digest=sha256:6b39fac84a792df8490da990ec7f2a03b6296113f6586838d4bd1c6f8a0af27d

Observation 84b7865a-05ed-4ac7-96ad-f9356ecc90d4 · inbound

Efficient Process Reward Modeling via Contrastive Mutual Information cites this paper.

Efficient Process Reward Modeling via Contrastive Mutual Information AlphaMath Almost Zero: Process Supervision without Process

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:25:59.087302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:40:04.180754Z digest=sha256:eae4c800e1c1c8484cdd905a7834b777f24e167c07bbea025da5ff46adc2c61f

Observation e1732b0c-e87f-47b7-95d4-6a4c3984777f · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces AlphaMath Almost Zero: Process Supervision without Process

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.970743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:870a6a64e2843f54881337a9436f3b8019642dfa93e284163bec589d6fff0d66

Observation f0f900f5-924b-43e0-9611-b935c93847ef · inbound

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models cites this paper.

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models AlphaMath Almost Zero: Process Supervision without Process

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:17:41.795472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T12:47:35.314486Z digest=sha256:64b34acd79abe479f2dfbe40217bdf242fa444121dc33cd8d16b76953967a545

Observation fc201724-9778-4f9a-ba21-49f95a24563e · inbound

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents cites this paper.

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents AlphaMath Almost Zero: Process Supervision without Process

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:59:38.180175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T13:56:51.914966Z digest=sha256:a4e13618a6bf813d8a305ef645396cd1a0400b7448ff0a279696c46c9d12daff