Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2504.12328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12328 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.601769Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:46:41.076602Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e82255d5-fafd-40d0-857b-3536af06b0d7 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601769Z digest=sha256:f72bc2934fb5e4ab2990c0872ef6afd3d9542c2f828e3843a512ee687074b142

Observation 7647961d-c5e5-4c55-a8a5-d33959af433d · inbound

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence cites this paper.

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:09.508714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:09.508714Z digest=sha256:8c5c997a8e731b1df22df45294b7b3472e2244a88f01b5320bb79e3016eb9cad

Observation 31047f64-36d7-4c8e-8248-7f2d24547fe6 · inbound

RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning cites this paper.

RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:33.906214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:14:33.906214Z digest=sha256:08ebc97d2158cf2a437648efadb8da6f3e1c736d2c22416678c66b452f400db2

Observation 15327041-fc96-46a8-b3bf-afecdf26e69a · inbound

Uncertainty-aware Reward Design Process cites this paper.

Uncertainty-aware Reward Design Process A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:15.150337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:15.150337Z digest=sha256:28109f9984b5a451e7ab730183049fea7016ec519a84cd3fb4449e690ee4dd61

Observation 44f79316-d684-4b68-affc-9868227b9b44 · inbound

The Other Mind: How Language Models Exhibit Human Temporal Cognition cites this paper.

The Other Mind: How Language Models Exhibit Human Temporal Cognition A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T15:30:38.748692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:30:38.748692Z digest=sha256:2851c3795a2de91f44cd07d3789a328f24a1bddee4ebe66bc7d6e7a94c4ea87b

Observation db472b39-e42e-4a2b-96a9-bfb7383bb47b · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:54.749256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:54.749256Z digest=sha256:dd29d0e57991c8ed6c0f7f670ec2e1817206b5d1878bead937fcba8fb83312a4

Observation 4f5f303e-8ec3-4e6a-8cec-2d4640f9a9f7 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.183065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:74e2c543ed991f7adcc35a9c70a9dadfa76f4f342b0367d428515da63432faf6

Observation 594a40de-242c-45e9-805c-50b82dea89d3 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 253

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:48.384610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:48.384610Z digest=sha256:ff91b4d3ceaeef14426369155bda56b12a3fff5143c7dce5c345c2eb5883cd0a

Observation 91b27b97-6836-4701-b574-f0eba351c387 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:45.216822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:45.216822Z digest=sha256:f4521457a87fcbc5ee782a830881ca7cfcc10dcfae088a9294d34d134708aa42

Observation c9f226c9-a0cd-4f28-9a2f-94e4b5493b28 · inbound

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization cites this paper.

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T05:56:17.734874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:56:17.734874Z digest=sha256:88d0e701de6878a4aa3efee6262548c76357612bc95b85b482b03ca66dc467ca

Observation b80ce6e4-1d60-4892-af25-9ed48cca04e3 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.645525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:99392fc07d417e6dd81e6feda1fdfe55ce61026218e2422392e5b38c46075f47

Observation 35a22826-9bed-4c54-ad0b-4b26450c99e0 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.230952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:f628cb207c175c2c0324683d2b3fc77b9d99957590e9b991074700b5a02b3b2e

Observation 39a11b7d-857e-4c3a-8c55-237c94b4d1fd · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.720310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:383577d3a3202b95d910d3313026c3609e2c4772006d621d781380a74296197f

Observation f89c9da0-0fbe-4a96-a8a7-c30ff9d84c95 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.637981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:67ce2bcdd2bad788268e534425ef8913762849e3372c083d4295c9161e166e0e

Observation 900a3756-db76-42c6-9c21-a578f0cea061 · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.393010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:834f02accd6d787d4d564e469f2ce1e32ef6496ad6499e2da0a5cb34a34a428e

Observation 61303caa-455b-4c9e-92bd-11e43e669e5c · inbound

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice cites this paper.

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:46:41.078116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T08:04:17.224987Z digest=sha256:d09d5d10cac10bb825d5c4e2663cc46f8d13e3b2870362b61c8c8ce3a000403d

Observation 0b1b94c1-dcdb-4cd0-8d61-8945a6a2d2d5 · inbound

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents cites this paper.

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T11:04:44.592375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:04:44.592375Z digest=sha256:0f7e1ef878be90039b42b9736581d8231a8b9ae40f79d3184d51edff6982e92f