Pith. sign in

Paper Citation Record · LEDGER

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 5 inbound Pith citation observations for arXiv:2505.14391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14391 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:38:50.332584Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:55.155806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:47:28.400784Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56ea353b-2cf6-437e-a849-016d3e2fb7b9 · outbound

This paper cites online" 'onlinestring :=.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.491408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.491408Z digest=sha256:307021f5288723dc84f1d67dd66c6704ad9b46f34038412aef901de09ccdf763

Observation 5ba3d376-3125-4cb4-bef0-b951c523e66b · outbound

This paper cites write newline.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.646682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.646682Z digest=sha256:cff1ee12417113988a71b7e183524b41bd587561835bad6f84e636878f1eaab6

Observation 99ca64c4-e157-473c-9c38-c563923aad0b · outbound

This paper cites GPT-4 Technical Report.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.810278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.810278Z digest=sha256:0dc44b9352f4cb822eca05a611a55fee4e34d49ddc8c59c8cc7ffb9c35eb0db2

Observation e6f6a7e9-f55a-4881-9803-c1929b3d812e · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.934365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.934365Z digest=sha256:a37bf43dbced70346f4e7ce1ce079c55a4db9907a324d77f1933068b012082d9

Observation 2913f427-467c-4a51-9ead-0fd50fbf8d14 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.053508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.053508Z digest=sha256:19631ee5c22252b5d6d0032ab9cbb2a6c752d1f198aa181bc5bdb5d6f1011a85

Observation 60664621-518a-454d-9e7e-1498e3ce7270 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.198208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.198208Z digest=sha256:00e9cc3295d1491cd63df60add74d11b04e661e79c1af0a9b6d0914ed08fbdd3

Observation a605f1c9-900b-4e70-878d-96b1991697b1 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.337373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.337373Z digest=sha256:b3d9a174651b0f0c82c524a505a53e65b6207b4d937e9d35ca69ec79b4035452

Observation 06c8340c-dd92-4d0b-98e8-fbecc9948eeb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.503915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.503915Z digest=sha256:fbc717d8a04eff3da66e398922991366533f96424ea5fb514994e4735912c880

Observation c1bbf3ff-8196-4307-b5f3-6b2c0075e88a · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.642136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.642136Z digest=sha256:04502d37abcbaf65ba52cd09e2684c6721602eddc9d47e31f882cff8654c9f2b

Observation 16e8b4a0-21bc-41e3-a6cc-843eb4a2c98c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.750931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.750931Z digest=sha256:5a75d149ec28f704ca1324560c605ccb82bb78cbd36fe378c6ed10d28aea30b2

Observation 1dbca456-136a-40dc-9da3-d3c8609eb920 · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.889283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.889283Z digest=sha256:e8968186dd012e3cc2d4a994419f4e5ed27ec0fd3db8b195b0eb20804b7f9dee

Observation 6a2055b3-5ad8-4212-9f18-dd114fb85d6a · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.040946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.040946Z digest=sha256:707557f56b6e35fdcf932e8b1bdc06939038b6757e2d66bc4109bc453419bea7

Observation 4f78a47e-f84a-49b0-a976-e0dae7057545 · outbound

This paper cites Let's Verify Step by Step.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.152880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.152880Z digest=sha256:e9d8ea19526e2aec40649663def1a2f26e36ea44e42d25b3da22bc078d096211

Observation eb3f3a78-6ed7-4ae3-8649-6891effedeb9 · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Augmenting Math Word Problems via Iterative Question Composing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.312631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.312631Z digest=sha256:bf7597c700ff9e48ecd8de7099f10f53d3dc350146f83f0013968b9dee9ad740

Observation 8a944585-18f4-4afa-b251-22318ae3ad1c · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.463728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.463728Z digest=sha256:3a6a1cd8ca629d6639dc0c920402921dd71685aede506207290e926890eb04f6

Observation 41a83954-97b4-4a1c-b509-692ce2a2caf9 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.611667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.611667Z digest=sha256:47ca181e24f15bf8b9c91c08d642400da60c46fd8c31e296a8cc666485343d1a

Observation f06f3f27-1078-4c93-ab80-8998c3a269fd · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.846806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T15:38:47.742752Z digest=sha256:e62de50f7c355f74dc3d6634dfa23a51acec86cdf2ba039b7820ab198793c013

Observation c24a8b09-f2c4-4ca0-bcaf-e6991b214e00 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.858019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.858019Z digest=sha256:06243a8ee32fbd33fed189034167a77c32aceaf21095380192acb909d1e267ff

Observation 5b37e09c-c8cb-43dd-b03e-b0f7df8d9fea · outbound

This paper cites s1: Simple test-time scaling.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning s1: Simple test-time scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.943305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.943305Z digest=sha256:361db96d3ac9d2aa98c4b87adf734eb0eddd2d1b14067f61f1b0322b0f90ff25

Observation 3acdf7a3-f627-4f9a-9956-ffd8bdd99dbb · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.144397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.144397Z digest=sha256:8894458b69548441c5c813f47f2d2e45c8a00debeffa672e384ba33dcf083e51

Observation c5437f33-0c61-4f58-ad92-432aa0b5a7c4 · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.712051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T15:38:48.277073Z digest=sha256:48ac96b0689be6b421f45ba4b419e355de9882cc6450077a633458ba0fd9a432

Observation 326fe5bc-3c46-494f-8c9c-400f7d0fc181 · outbound

This paper cites Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.422552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.422552Z digest=sha256:8f30df1e03dba949d44a098019f37afa41264360d933875ccc7b807f34c01fe9

Observation e003c48c-8e24-486b-a8ce-5a086ee337bd · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.565575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.565575Z digest=sha256:4e14274161c973291f800b170e1751aede07538e7d44e848ca79719370598de4

Observation 3a7b88cd-be1a-4b3a-bf1a-3897f50274cc · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.738925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.738925Z digest=sha256:db47725b45c0440e5c708d9dbd825fbbbbf6f8360816d568ab066f14820ad22e

Observation fc93cedd-75b4-4182-8da9-3dd5f9219ee9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.946473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.946473Z digest=sha256:53146bb3bb16d64d2288afbc6b08e0bad6a570f859b127a5e6cf30e21d84ff2a

Observation ca3e96b9-940b-4df8-8af2-26d93ae340af · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.448514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T15:38:49.158853Z digest=sha256:b8e71568af5a85cefd21c7a541a70d5ad4fbff396570ea6f5c1dda3cb8b4f3d5

Observation a30162a0-c2bd-4ab1-9c0c-bbce5076f09d · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.195028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T15:38:49.306606Z digest=sha256:6b418febb0e629b9573e47e967e3952e27bd20f55eaa06abbaab6c2ea5d5d838

Observation 5eab2423-d5b3-4a0d-be5b-c291cdbcc078 · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.425108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.425108Z digest=sha256:0df656f00a01a154b245e9641426253548a7b87f37afd7ea333a161472d06ed3

Observation abf49cd1-0d55-431d-8f3f-9634012cd3ee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.624831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.624831Z digest=sha256:b707a5a3a96558bd4371e047b8e88c86efa34f444974509c641a25be81cba4d9

Observation fbafd562-03d0-411e-8ae2-bf6c467f52eb · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Solving math word problems with process- and outcome-based feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.778223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.778223Z digest=sha256:b8080b7717fb7d0f01bfa17f21494ed3938d4e7b4069a31840bfa01dc23f8afa

Observation 64034c53-e4af-46bc-87c3-a756aa20b3df · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:50.970290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T15:38:49.897802Z digest=sha256:c58e9edd3983eccc2433711ae6452565c75bde336ab7f75994eb37039e419bf2

Observation c9c70a28-981a-443c-b618-94a132a8e05b · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.957959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.957959Z digest=sha256:f86d13270438db913499b48d78b5fd4157de506735e0b83da2d21b1d69ee0979

Observation 4d06186d-a39e-4144-ba60-7034a482949e · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.036904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.036904Z digest=sha256:6008e02be572b98d0aa3efe662f1c0a57f53c616ec77b440acb1ffbc176512b4

Observation 5ab26c4c-3faf-40a7-a2b6-f62068ae0116 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.099971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.099971Z digest=sha256:6a5c9a32ed6874e1194665bfa1b01f95b1b5cd8bfafe405d29531bab3e3f79b5

Observation 53065805-6222-4cb9-827c-c369526c799f · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.179847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.179847Z digest=sha256:0d6b89dc723b968d8b3df0a8055a60eeead087e807acc7cb0e19f3c2b25d732f

Observation 630a15a5-4fd0-4948-9489-2af1c8d60d94 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.251756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.251756Z digest=sha256:c57ac69fe59ab73999ba5cc81eb4954747ff9d5a91781d73e1db0ecf45ec63d5

Observation 32fdfd0e-e858-4043-823d-a09a5e0bef91 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.332584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.332584Z digest=sha256:16b3f6d2aded8fe4d27ce7f7e37190a40d04cd5998368cdfde73a21db6c4307e

Pith citing papers

Observation 4bdfda50-1c54-4dda-9721-25b4f9f42947 · inbound

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning cites this paper.

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:55.155806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:32:55.155806Z digest=sha256:2989a632fbeb2dad3e321ac9496dac2c8ecd123425f99af05508712ebd358d2b

Observation 7e729a3c-92e4-4d7c-a159-81f0384848bf · inbound

LLM Reasoning with Process Rewards for Outcome-Guided Steps cites this paper.

LLM Reasoning with Process Rewards for Outcome-Guided Steps Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:47:28.403595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T06:43:19.869111Z digest=sha256:2d2590abc6882174551fcc53bdce8da0dac9c7587d4ed7de3a9222b1f11965b6

Observation fe7a82cc-eca3-4c9f-bacd-f3766a158e7b · inbound

Improving Medical VQA through Trajectory-Aware Process Supervision cites this paper.

Improving Medical VQA through Trajectory-Aware Process Supervision Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:06:06.051115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:18:17.846140Z digest=sha256:6c8121765684a5d37ce5e83459e1cec1aefa9f2367ab725bd4235144307ad34d

Observation 16ffdba1-42f0-4c17-8c90-bd1ee5a4ba35 · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.809955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:3999d3ddc463ee7c3841dfe003527f05a9a5c8af65b75c78e3ba1b54c1851aee

Observation e49ff7ad-a288-4b99-92e4-8b31bf123d1c · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:22:50.477238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:65a4c3105e63a435a99cd0c90fb605a4ee262b2e6cc664f36908c5e5b5627b5a