Pith. sign in

Paper Citation Record · LEDGER

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

As of 19 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2505.19706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19706 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.519840Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:58:49.973078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:53:06.955702Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c6f086-183d-422b-a08b-8f31bb64ebc8 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:18.968255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:18.968255Z digest=sha256:9711b28e39e2846b0791a0e8947159df190240f88f1ad1f65b7ec72ebf8b2b62

Observation 6d751b58-6292-4e64-b21a-6bb4a7c6ffe9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.156320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.156320Z digest=sha256:b933d0dcce53cf1a7d5cd416869b87c0196fa41f3dc48b1428c4c1ec80da2f12

Observation 5011179a-17b5-44e5-85fa-4f9aeb89f488 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.237664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.237664Z digest=sha256:9d42a3435ad2d2032f48ac6a553d58b3c94f51075baa2e621b134f8e68ed0078

Observation 2adfa8dd-79b8-4883-979a-77031df66ba4 · outbound

This paper cites Let's Verify Step by Step.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Let's Verify Step by Step

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.319766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.319766Z digest=sha256:7d1cfd5b45173329e88128368c95e81d998c855d3a3f037b8291187aeb99dbc2

Observation c9834023-d4e3-4163-b2fb-b9b35d000c63 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.428465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.428465Z digest=sha256:68d94c6c10937eb41f1baf7c9b19e647a7ef5d080b3234816914ed21addd5df3

Observation efd378be-ec6a-4293-9a0f-93cb4fb7e8f6 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.554437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.554437Z digest=sha256:e6fd17d2e69a8dcae0ee820043a3452a69d5a549e6deed4459bf68e834666d32

Observation 8513cbaa-decc-4914-953b-2c653543a98a · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:22.011172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:14:19.674351Z digest=sha256:01e6b9c50942f62333c205e375f68cb322e57d93dc525c9fe959d9c7ea240e67

Observation b66f5c06-a081-477e-87d8-c6ef3911f85e · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:21.835635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T14:14:19.783614Z digest=sha256:05723780c21bdfe9da84e0f804143e1095ecc1497f171cc0d0011e4adb92094f

Observation a86bc2bd-c4a8-4c07-8e90-39319c29e13c · outbound

This paper cites R-PRM: Reasoning-Driven Process Reward Modeling.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision R-PRM: Reasoning-Driven Process Reward Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.876368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.876368Z digest=sha256:3a0c69805cd158314a259fd88725905248d82ba5958c659bfb908977fab49af8

Observation 58760b42-ef1c-4624-b54b-aa01c9a01cad · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.989779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.989779Z digest=sha256:f8126fc99ed1b0ff3bb85bb8ae0070547e20d12312292e70c027b713c6c62a1c

Observation 56a5b2ec-712a-4545-8dd9-8ba54ba819fc · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.136880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.136880Z digest=sha256:d639a71e0322bcb748330e3abe15139f82f33e272ff410aa0c43148a1284f79a

Observation 88c8303b-0f73-4d43-88c7-98bc4cfb0d3d · outbound

This paper cites AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.254758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.254758Z digest=sha256:c16693944378ee59407048cc6d021cd58371161859848b055a198da06b2cbbd5

Observation b52878ae-1e85-4ec6-9057-86ea82c1a8c9 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Solving math word problems with process- and outcome-based feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.364504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.364504Z digest=sha256:72ff2b7a8ede076b0b988a6994683f512ad48f0132073800738c0d982bf6f9dd

Observation 879e5c96-25ab-4a97-b84f-7b061a93b65f · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.522619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.522619Z digest=sha256:1a37da75d1ed20bed74e3f83432723d3a9cb934fbc336e84adb84135989af99b

Observation 89155cee-1c19-481b-823d-169683158bd1 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.610933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.610933Z digest=sha256:93ff6a25fd4bd80f5be8f5f0faa528f0b3b93948a8601e9cf0756e9af5918830

Observation 55b9c388-e4ae-40c2-8c1a-b682b88306f3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.694287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.694287Z digest=sha256:2e42f64008f3517518de17f2e3be13d214defb5525e7a5d559efcc8a5caa0732

Observation c640f3d6-9084-40a9-a532-b9e2b78e8fd3 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.773480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.773480Z digest=sha256:1bb45acf34860f7f2afbbf885c784425512c241e4289bdea1ee7444ac89a7de9

Observation cd0f8c5f-24aa-41c6-a3e4-cdccf5033517 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.861017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.861017Z digest=sha256:d58ecbfac8d64e53d8be41cc57d85308640578367efb4e167de3e11f110115c9

Observation ba321121-99fe-4196-8d6d-d1a4d1c5fce0 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.931464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.931464Z digest=sha256:66ca06ce07aaf7bd454cbac9302fb40c229068a2860d6d8edc98569587ff9728

Observation 828d72b4-303a-42e1-bcc8-3f04d1b990c8 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.023530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.023530Z digest=sha256:7d71094489812c7a604a4b8bf8c9b3f20486ecbd3b1f16f9af9a25b01cdd161b

Observation 9120d802-a4f3-4240-ba10-d949d515721a · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.092331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.092331Z digest=sha256:9862ce6ee7640c0c9fff73089128cf0d5cf17bcb8161c5df9a45c54de460a073

Observation 85eae2ab-d0f4-483b-9e3f-d46c1d24eeb0 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.168455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.168455Z digest=sha256:ed326ea27ca572d64bbdd3839365fe497106007a958581b907cec4f08c719ace

Observation 0606978e-3802-4f8d-be98-70a2a53d2731 · outbound

This paper cites GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.245461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.245461Z digest=sha256:d719939409ccba0a5e07ce4b37e0099cc7b28f8e8e485f44bf21cd41d47e95e2

Observation 54282e4b-965d-4c47-8922-efafabd17d13 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.314240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.314240Z digest=sha256:d166d45bc0f2b3f36f4a5a7c88ff49b1281afc8613456b68e6569663d2ec553a

Observation c8b28158-9ba8-4576-bf51-3fbe86659bd2 · outbound

This paper cites online" 'onlinestring :=.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision online" 'onlinestring :=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.388599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.388599Z digest=sha256:3462725f6f4cc3f8afa09a9a7ffe28a62e5bc3170efd1a3cf6215015b5f3119a

Observation c2acf720-a52c-4464-9441-157725fd18b0 · outbound

This paper cites write newline.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision write newline

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.519840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.519840Z digest=sha256:c4f03bd69428c6d0c1c6668426722e4a3f0ed5ae3ba11ed6c730ad680a9664ae

Pith citing papers

Observation a1b42452-0fdc-42ee-af57-006b3dd497e1 · inbound

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models cites this paper.

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:49.973078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:49.973078Z digest=sha256:a4e0e330136049a4efe6315cdd7a35d37ae74a7434792a202823122b9a8e4888

Observation bc392757-c65a-47c2-ae0a-5adced636e00 · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.957200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:6f02a1be9934bd024da5ace2cc4c770a0e2b197e5cff26dcde3688768f0ef029