Pith. sign in

Paper Citation Record · LEDGER

PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2501.03124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03124 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:40:36.052680Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:37:30.074521Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f3cb51d7-1b83-495d-bf7c-ca1282b5ae79 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:33.020711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:0a88067080176a478b7692c22a35c79b59ab36764c6650d92d1b64d6f9926a29

Observation 37177012-0210-4d98-88e6-f2e2220604b7 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:36.052680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:36.052680Z digest=sha256:ff1ec17dffddc640bf6e0c92203596b262e8ffd721e0997c9214d92900405a6a

Observation c5c0f5c7-b262-492c-a499-18a3a9c475af · inbound

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts cites this paper.

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:39.822263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:39.822263Z digest=sha256:5609e89f4b3b76e77762d287a2e33b8791200ff105728b807c3b1c909ee8aa40

Observation 58760b42-ef1c-4624-b54b-aa01c9a01cad · inbound

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision cites this paper.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.989779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.989779Z digest=sha256:1e0d67fa7be57d67c657f6a583eef59e90b3790709adc1dbed78c4d4e478502a

Observation c4fe288e-8194-4497-8688-e4265491cd41 · inbound

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving cites this paper.

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:58.204424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:48:58.204424Z digest=sha256:03755d0d4dda0520b4259169134716a3a6b417a88a460d2dd2859d4c5e88f9ae

Observation 495b41d2-2427-436a-9c92-bc82e99c77d2 · inbound

PixelThink: Towards Efficient Chain-of-Pixel Reasoning cites this paper.

PixelThink: Towards Efficient Chain-of-Pixel Reasoning PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.805388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:45.805388Z digest=sha256:3243a5c69644f364efce60624578f64e629b24c6985e42b5a8b68d5ba74518f7

Observation 22ae8d7f-691c-417f-975f-b2e50f5f1b5c · inbound

Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively cites this paper.

Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:45.058024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:11:45.058024Z digest=sha256:517f84275bd637a4deea7b2faa7e64936d4963203b8e1a4c32c64f4786f5787a

Observation 1497c494-db3c-49fd-a62e-91d9c355b02a · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.761375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:54a6b3462ff10750b86aa96c6bf424f5a17b68e77e10f6c011fb67dc8dad5262

Observation af15ffc1-4b55-4d7b-893f-55462028cbe3 · inbound

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models cites this paper.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.154021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.154021Z digest=sha256:52ef030c3d3c07f4716b0657ee37c78b9d5803a9cc882cfb35ba165c34356806

Observation cf571c63-c71a-4caa-946a-eea22154689b · inbound

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models cites this paper.

Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:24.928392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:37:24.928392Z digest=sha256:cb54200ef1d4650d499b8e06577fff344c491e651ccfe3f076304409a8a11f58

Observation e5dbde34-2e51-4319-ae54-75709c18902c · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:53.215356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:53.215356Z digest=sha256:49b86db3c0649c2cc0ba59d96691e7092c9e28a59603546d1f7c47cb5dbfe8c7

Observation 8405bc8a-abd0-4739-b46f-0993e3b2409e · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.513775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.513775Z digest=sha256:e4005d2b54301ad488fbcb08043da809cc91fa1b90018b5c39c89aec00d0656d

Observation e0312c8b-6a8a-4d40-8800-37cb59664f0f · inbound

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models cites this paper.

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:50.210843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:50.210843Z digest=sha256:8ea3db21dcfe1470106aa335072ee097950ed26eeca99b2ee1e43bff3f9d2630

Observation 7464e93f-84fc-44e9-a514-0b3aa0c65612 · inbound

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models cites this paper.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.114132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.114132Z digest=sha256:c82fffbea782b08d147f7bbce5832c1649604bf9b8fb8b65df6215ccb478f143

Observation f5835ae8-14e8-4b42-88d8-b6c4aca44a57 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.374832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.374832Z digest=sha256:a04a7ef2eb62f45caf39d45ca71b1b853c3540c966cd168f9cc765676c4caa2d

Observation 86e8c7d4-5a21-4c29-9cdc-f35c797c9d59 · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.372438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:0edb27fe2794ac097b8f20f4f57a4facbeede47719b14c710899087d1455b93a

Observation 9e70fa15-b526-49f0-a0aa-f0678415e98a · inbound

Self-evolving LLM agents with in-distribution Optimization cites this paper.

Self-evolving LLM agents with in-distribution Optimization PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:57:09.539333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:18:27.021136Z digest=sha256:fb99269a68f670f34838fa344f03303c580cff6edbac9bb71a6b2c5c7327b4ab

Observation 1df5a118-8c41-40d7-8d8d-60c56b63667b · inbound

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning cites this paper.

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:37:30.076523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:07:06.227106Z digest=sha256:414d0b8db5f28353c06bd81e9c4a62605a5620ac2c6e28b2be05b7707c70f7dd