Pith. sign in

REVIEW 1 cited by

An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02382 v2 pith:BIT5NWRO submitted 2025-03-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningepic50kmodelsprocesstrainingannotationdataepicprm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) is of great scientific and practical significance. Researchers typically employ process-supervised reward models (PRMs) to guide the reasoning process, effectively improving the models' reasoning abilities. However, existing methods for constructing process supervision training data, such as manual annotation and per-step Monte Carlo estimation, are often costly or suffer from poor quality. To address these challenges, this paper introduces a framework called EpicPRM, which annotates each intermediate reasoning step based on its quantified contribution and uses an adaptive binary search algorithm to enhance both annotation precision and efficiency. Using this approach, we efficiently construct a high-quality process supervision training dataset named Epic50k, consisting of 50k annotated intermediate steps. Compared to other publicly available datasets, the PRM trained on Epic50k demonstrates significantly superior performance. Getting Epic50k at https://github.com/xiaolizh1/EpicPRM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

    cs.AI 2026-07 reject novelty 6.0 of 10

    Directly editing a large reasoning model's chain-of-thought and feeding back a distilled version of the edit improves correction success by over 25% and cuts token usage by roughly 40% versus dialogue-based correction.

Pith tools