Pith. sign in

Paper Citation Record · LEDGER

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2505.19706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19706 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.519840Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:58:49.973078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:53:06.955702Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c6f086-183d-422b-a08b-8f31bb64ebc8 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:18.968255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:18.968255Z digest=sha256:508504a5ba5f1eabc7e0af9bc6ce8cda0a5db8c5ae658a990ee9dcae42ba8c0f

Observation 6d751b58-6292-4e64-b21a-6bb4a7c6ffe9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.156320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.156320Z digest=sha256:550677d7ae8dc9e3f2971b8101d2e3f034fa742f2e1a2e7b19ab6276cc440ace

Observation 5011179a-17b5-44e5-85fa-4f9aeb89f488 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.237664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.237664Z digest=sha256:7d65ade2e63d696ebebaaa36fa7c0d5bff8ba42e1bee1ac2b07bcf2faef2ed34

Observation 2adfa8dd-79b8-4883-979a-77031df66ba4 · outbound

This paper cites Let's Verify Step by Step.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Let's Verify Step by Step

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.319766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.319766Z digest=sha256:075fb406cecc58d833ffd4f493bf27ac9d5b158b49a2656d17bdf1bb6d71c977

Observation c9834023-d4e3-4163-b2fb-b9b35d000c63 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.428465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.428465Z digest=sha256:d5594764a4a2a059c217773251fb160c5c53b8ec95b838a4a411d3b96ac4b68a

Observation efd378be-ec6a-4293-9a0f-93cb4fb7e8f6 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.554437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.554437Z digest=sha256:bbfa299e455f968ba614b092e7e28d2477b0243971fa0e1d5514e6f775adf84d

Observation 8513cbaa-decc-4914-953b-2c653543a98a · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:22.011172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:14:19.674351Z digest=sha256:303f34c763078758c2f42df93563b5a14f8f92b354a9dbe985465650f454f669

Observation b66f5c06-a081-477e-87d8-c6ef3911f85e · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:21.835635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:14:19.783614Z digest=sha256:2ca2d8a6a1743484d77830d858cfccf26f221754ca1e3bf5a7af8d01ab5e2439

Observation a86bc2bd-c4a8-4c07-8e90-39319c29e13c · outbound

This paper cites R-PRM: Reasoning-Driven Process Reward Modeling.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision R-PRM: Reasoning-Driven Process Reward Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.876368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.876368Z digest=sha256:3a0f7a8070028f23baaf272ac7ffbe0302abf5adeaf6556f8fdb4762273eeb41

Observation 58760b42-ef1c-4624-b54b-aa01c9a01cad · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.989779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.989779Z digest=sha256:1e0d67fa7be57d67c657f6a583eef59e90b3790709adc1dbed78c4d4e478502a

Observation 56a5b2ec-712a-4545-8dd9-8ba54ba819fc · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.136880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.136880Z digest=sha256:12e8340eacae8bc3931bb8c0ee6e596b24cfad4216357df3db16c99680854aa9

Observation 88c8303b-0f73-4d43-88c7-98bc4cfb0d3d · outbound

This paper cites AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.254758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.254758Z digest=sha256:caa66f564f66fe02197945a8b728ae8fd804f0b5fc8e985daf586031fc91e6a0

Observation b52878ae-1e85-4ec6-9057-86ea82c1a8c9 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Solving math word problems with process- and outcome-based feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.364504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.364504Z digest=sha256:b7ae86c6592b0ddf4d8c35603281e2743024ea85eef27f4e06ad0e26fb8c3667

Observation 879e5c96-25ab-4a97-b84f-7b061a93b65f · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.522619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.522619Z digest=sha256:2ca7d26dd1b696c3a726c98122aee95aa58fd4bddcacf21dae018180266d68e7

Observation 89155cee-1c19-481b-823d-169683158bd1 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.610933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.610933Z digest=sha256:b18a9dc8d1bd20194f5ff632892f7bfbdb51df5882bcae9fa401a133b53afbd1

Observation 55b9c388-e4ae-40c2-8c1a-b682b88306f3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.694287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.694287Z digest=sha256:467560c86780a9368e1ed9030317a03d152046415f946d195b0b0a314a806fa8

Observation c640f3d6-9084-40a9-a532-b9e2b78e8fd3 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.773480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.773480Z digest=sha256:ffb46b42c90cfd39cf3e6655807d4c5d6392525a5a72f904fa93a2597cb2a4d5

Observation cd0f8c5f-24aa-41c6-a3e4-cdccf5033517 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.861017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.861017Z digest=sha256:864386924c7f9b73ea77ad500f0439fbbcc77c5cbe6b4e9687591f44a1762acf

Observation ba321121-99fe-4196-8d6d-d1a4d1c5fce0 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.931464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.931464Z digest=sha256:7628f3c2307f7922d4495b384c2610b0bfb4695bf084c3da7ff937510e114fc1

Observation 828d72b4-303a-42e1-bcc8-3f04d1b990c8 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.023530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.023530Z digest=sha256:e15da558b41400c4b43d50f25cddd81dae22a9b949aeb86c5e2c587263758f4a

Observation 9120d802-a4f3-4240-ba10-d949d515721a · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.092331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.092331Z digest=sha256:91024c5d65ae5dd1f656f3c86824d07d72ca7ad4f4a12bb61dc2772863280189

Observation 85eae2ab-d0f4-483b-9e3f-d46c1d24eeb0 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.168455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.168455Z digest=sha256:b32b2fa000a5aba85185a0445c85708859519eaa72a4343927ea1d0892f352fa

Observation 0606978e-3802-4f8d-be98-70a2a53d2731 · outbound

This paper cites GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.245461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.245461Z digest=sha256:15d7bcd0fb5e09bbfb821282bd2ec8c81b5f50837d7e0253f0cbf6499e4de562

Observation 54282e4b-965d-4c47-8922-efafabd17d13 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.314240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.314240Z digest=sha256:ee5ad15c0469a4dd7550b1a65ae5c1791572bc2ed599aaedf56807f2daa04b35

Observation c8b28158-9ba8-4576-bf51-3fbe86659bd2 · outbound

This paper cites online" 'onlinestring :=.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision online" 'onlinestring :=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.388599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.388599Z digest=sha256:7c717af5cfaa0cb41a9592b9fd6ebca6943bcae3ddf5bbb70c700359eeb56c77

Observation c2acf720-a52c-4464-9441-157725fd18b0 · outbound

This paper cites write newline.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision write newline

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.519840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.519840Z digest=sha256:d18693e995dd73fc3fa40836400252944c68404bf90f38b90129f0fbc1b12016

Pith citing papers

Observation a1b42452-0fdc-42ee-af57-006b3dd497e1 · inbound

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models cites this paper.

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:49.973078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:49.973078Z digest=sha256:2b82cc4ade27ba08f0ad5e3d8053926bcaca5a127598f08e539c6b83dfe97452

Observation bc392757-c65a-47c2-ae0a-5adced636e00 · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.957200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:5ddeca199f03867aa087461c150ab32bdf739079029212659300bf8dbd60e2fb