Pith. sign in

Paper Citation Record · LEDGER

FreePRM: Training Process Reward Models Without Ground Truth Process Labels

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 4 inbound Pith citation observations for arXiv:2506.03570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03570 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:42.841133Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:57:32.948217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.479928Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b692a8df-3cdc-4c89-be02-b60657262463 · outbound

This paper cites Alphamath almost zero: Process super- vision without process.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Alphamath almost zero: Process super- vision without process

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.644217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:40.848523Z digest=sha256:bf85e4c7cdb1dc201bc015dcb2f1157aed3f9c285dede766fa55a9ff8fe70444

Observation 1f6013f9-299f-4d41-8515-9aba3e27ce84 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:40.892903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:40.892903Z digest=sha256:7006bc1ce680c6fc0c251d6dbadfdddd9cc674e2e1d3c41e92d080c02427d5ed

Observation 409cec1e-f603-4b34-9b8c-778ef84edb9a · outbound

This paper cites Process Reinforcement through Implicit Rewards.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:40.984203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:40.984203Z digest=sha256:7f5001da0f1da33c39fe7c98374f961bfde395e0606d94f784e0c06974cb8339

Observation ce41dbb2-d935-40d8-917c-c8f0874ef52c · outbound

This paper cites The Llama 3 Herd of Models.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.062201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.062201Z digest=sha256:b7a919f597d1f7588427c6ed22dfc97926cc0bd6f566a6d99872282d75ebf9e4

Observation 60386761-00f8-4d14-9105-f477aa201bba · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.126923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.126923Z digest=sha256:503081a50d26d7cbc400a5b6dc182c71a4d121ef36b06abb9b890459ca8d1ee5

Observation ae69aec8-0373-439b-ab6c-ce25584e52db · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Measuring mathematical problem solving with the MATH dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.562843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:41.206238Z digest=sha256:3603d4cdc46a0d662e0e139898e5e654806ebd33319808660db3acc3f62b0499

Observation 07e9f2ba-67cb-4956-8cec-c8d6ef4c59e1 · outbound

This paper cites Qwen2.5-Coder Technical Report.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5-Coder Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.250937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.250937Z digest=sha256:ad03c4fa14b870da79892cc454551019e83ee120fe99969876073227c96f5708

Observation 5d6ddf36-bc9b-479f-b39c-69e0daab2ff2 · outbound

This paper cites Mistral 7B.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Mistral 7B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.343356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.343356Z digest=sha256:0d5fdf9269b8f9bae50b4d0a8b00048b76094765448621d23bc7b6d508855c08

Observation d879184b-5a0e-4540-b615-95c2dd7eec5d · outbound

This paper cites Process Reward Model with Q-Value Rankings.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reward Model with Q-Value Rankings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.397499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.397499Z digest=sha256:34534db6002234f6b25f73261de19618b0e46e63c9f8b5db29db01865809e1d4

Observation 41423e53-92bb-439f-b9d5-05f4325a7b4b · outbound

This paper cites Let’s verify step by step.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Let’s verify step by step

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.434667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:41.502360Z digest=sha256:9e20ed33a1ec0657bf39dd78cb48634ced17e9157996666c78c02fb1b702aa08

Observation 768ef0a6-eb07-4df4-ba05-4b74f7a63012 · outbound

This paper cites Autopsv: Automated process-supervised verifier.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Autopsv: Automated process-supervised verifier

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.321785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:41.560493Z digest=sha256:0d16a54d065ab7dc77e5d535fb1498440d21188d62fa5252fd3a8dd2ea17e80b

Observation 71f147f5-9355-44db-9ffd-cf487da4571e · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.654194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.654194Z digest=sha256:dbdb1d3546f1a6b9a7bba24f2649cb13abfe07839e97c9cdd4cd548415f27953

Observation 2107b35b-0ad3-45ea-b7eb-01f31268c4ff · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.718662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.718662Z digest=sha256:fe6ae09a27df23312cd95372d0347e27e57b8e8220101f0f81db250e0f87401e

Observation c138dd29-192a-467a-aada-eb6e5c7cabaa · outbound

This paper cites Skywork-o1 open series.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Skywork-o1 open series

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.213012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:41.762843Z digest=sha256:3a749e38dc0257a2b0cc53271091c7ad088cfb08adac48ca5f74131b10d5c446

Observation b94bc281-9988-4294-afbb-d713b7ea1435 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.911560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.911560Z digest=sha256:d5a757870d7fa187ed95242ac487f02d54c45221c80ab48e331c62851d8ebb2c

Observation e88e6dc7-a111-4a8e-8439-e4c2192d5cd3 · outbound

This paper cites Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.008548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.008548Z digest=sha256:97c13b94733f5cc7f726048c14c2455f8c1a8ddd781baed08e17cc1cf4c048d2

Observation 67cac0ea-15ea-4643-b7c7-f04d64faecf2 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.963514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.057605Z digest=sha256:8c48d6de08225cd96033092405175c1339f6906972ae7a42d32b00fdb6ee7e63

Observation 9a29aa4f-7abc-4e3f-89ef-5b847debc3a2 · outbound

This paper cites Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.838113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.140691Z digest=sha256:17e26e00d64b5595f98209420ada09ca21caecf061dd90e36ce6ab0f745028db

Observation 67831e97-c489-455b-9ff2-72921486e6ab · outbound

This paper cites Training large language models for reasoning through reverse curriculum reinforcement learning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Training large language models for reasoning through reverse curriculum reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.758576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.205649Z digest=sha256:e2116ac76afeb50e1910ba72cd1899510910bf7176a53a9c5fa31398ceb9f2a4

Observation e37d8d9a-6eea-411f-a8c9-0c9fc9a8503f · outbound

This paper cites Evaluating mathematical reasoning beyond accuracy.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Evaluating mathematical reasoning beyond accuracy

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.633494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.272829Z digest=sha256:6278ab2eb97acb957d4a14e71ef62f1ecb31fa8bde13e5b6898a742faba096d1

Observation f102bf8e-4392-41a8-91f9-5159c003ec7f · outbound

This paper cites An implementation of generative prm.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels An implementation of generative prm

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.509443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.339628Z digest=sha256:ed0d8db79898b92942fd59586bd6db47337e73a3e241b601ff8aad419f744141

Observation cbc7beb3-2e64-4474-9d5b-61baa3ed4d88 · outbound

This paper cites Qwen2.5 Technical Report.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.395649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.395649Z digest=sha256:6824e5e7213ffecafc9382f190cf5761cb773adae9b2aeb88ab53ea2dc386c5b

Observation a586e8ac-6143-45c3-a3dc-3f523d1d4f41 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.449307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.449307Z digest=sha256:01221ece5489b318dabee4c17704b72773aea636c33a8ebefe3a56bd673f74d8

Observation 64cba8e6-cf3c-4a15-a310-e833bd8a7636 · outbound

This paper cites Ovm, outcome-supervised value models for planning in mathematical reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Ovm, outcome-supervised value models for planning in mathematical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.348491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.539085Z digest=sha256:ae728139d65d2e6a42347e0115a20ecef607a64a6c16e7c799c1df81910327e2

Observation aef94d07-53cf-45ff-8dfe-8b2e80688aae · outbound

This paper cites Free Process Rewards without Process Labels.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Free Process Rewards without Process Labels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.575866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.575866Z digest=sha256:c0768eae184351cd7c8488bc080d589b046dfac08edd47c659eee00c8c6c5916

Observation 944cfdcc-93f7-47b8-82ea-1109dd1f95d8 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.642734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.642734Z digest=sha256:e809f6c4da7126e0d4bacd27efffb57dab14595570e0044c4c3631778e99bc2c

Observation 814b695d-53dc-4b7b-b30c-487d35b11c74 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.722080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.722080Z digest=sha256:83635ee8a8d2b8a0462494b741dcfe576cdb9d9d6c7fb23d2fb2a1064b9dceb2

Observation 63478cf3-b4f8-4893-85c5-e0d097331ed0 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:06:42.769583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.769583Z digest=sha256:f50d1da2360bb049397133e8a8d4780c510198dc1cf47629bc3de1cfd35dd438

Observation 62e255ff-a0d4-48b2-8583-4ac3d54511ee · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.161376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:42.841133Z digest=sha256:a15cb0cc825167a0fc056eb7824578db9a174df91ae8538491c92ee217f7c70c

Observation 8e862179-7b9f-476a-ab48-33c92b429597 · outbound

This paper cites an unresolved cited work.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.086411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:06:41.854876Z digest=sha256:6b1ddc14dcedb83916c86d5ced0d27bd7332499491b8a034a904d01d586e0bf4

Pith citing papers

Observation 355ce1e2-f6ca-4d58-983a-4b788d6b6edb · inbound

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment cites this paper.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T19:57:32.948217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:57:32.948217Z digest=sha256:e1f72e5c06ad58994e75b01c916e92f1fccee7917ca112414895a7245fbee9d6

Observation d1a55f1a-dc7c-49ae-9cea-5bfb614c2268 · inbound

Unsupervised Process Reward Models cites this paper.

Unsupervised Process Reward Models FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:18.892131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:19:04.275069Z digest=sha256:452aade82a8743af3cf4f36b451d7ea94c22adb4c873f7df0a076f276723db67

Observation 2dc7b6a4-248e-44e1-b692-e4fa63a8c41b · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.933322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:9dfe40908fb901cfcbd6adde7e32f306ec3180e959164c955cc0a6d049af4d15

Observation 32756f27-4328-4fa0-ac5b-0dd79cc3f4d8 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.481368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:464e7c7ee71f89f8403d13bd5af2212f3672ad3f94227c5aa16bd3783203ca80