Pith. sign in

Paper Citation Record · LEDGER

RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.09893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.09893 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:45.894072Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T06:06:43.201973Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d474fa82-b0e3-4186-8838-bca08a1cde27 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.848180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:31b33a807c042ee775746e2734cfc5b72bba35db2a40516e9291f77ba5a0957c

Observation 44a8b972-e68a-4d35-9937-36f6c758c184 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.747759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:e1c28902f85c71a7db53ece2f932f58072ecd966816a872c190a596d6931e10d

Observation be8ff1ed-8400-433b-aa1e-e13b8a79b760 · inbound

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents cites this paper.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.894072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.894072Z digest=sha256:44973b110f0fb64db1f4b63e3fb3bacbf80a8a63efd39eb4daabfb39a36a97ad

Observation 87686c92-4e7f-4a09-a47c-0fcb6213e753 · inbound

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards cites this paper.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.702068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.702068Z digest=sha256:3bfa5bdfedb705778c700c4fa138c9a845c10bb1bd14fe090616397ca987ad03

Observation 0d4aa580-82cb-4bd6-ab26-895ceb6edbdb · inbound

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models cites this paper.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.443699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.443699Z digest=sha256:9981533906f247423da1bf2a8aa4964f0873b7df1b9db2b91ff0eb9734ef377c

Observation 59d294e0-8a4e-4220-8669-a01adfbffc31 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:45.244103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:45.244103Z digest=sha256:34ad286fe204e976676c56a5575d7e44be7a863dcb60cee45d60b781dfcc48f4

Observation 637b1400-4ceb-4add-9383-c99153db5a94 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:50.350098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:50.350098Z digest=sha256:11931af723503741726823184f051bdbdf61b43f34af08d4f75d0cc2ff4b6772

Observation 552759f0-cecb-4c3b-b700-587d580dc87b · inbound

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization cites this paper.

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:18.350925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:28:17.333001Z digest=sha256:1497aca2276f43f2b2a8a27efdd8e27e2ae9fc13cb389dfe8787ea8532ee0d13

Observation 70897c33-a790-4525-8ca9-29f8cadad99f · inbound

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization cites this paper.

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T16:39:51.114742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:39:51.114742Z digest=sha256:ba3f1714e9af77d0eb4c249c82663d9a5d7f37a8175c5734c737e1d8c2196da1

Observation e5900649-fd9b-4461-a55e-64c5f27f6854 · inbound

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences cites this paper.

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:21:07.077834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T17:27:05.533755Z digest=sha256:9c7a201a6905511243cf2a4bd569cd3f50fa1507a3daa178bb11d9d7467734ce

Observation 69088732-1197-438e-a9d6-2f62f6f64f58 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.021813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:daefeb57005053647667b9f40d079cb0f62b85152c52a5073662311f31067c68

Observation 9b7a6ff5-e858-41a9-b9f5-3931a8b0bc39 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.204559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:fba9505b62e743fafb22e2fdff38a533662e2261fbb4ba2a85d2bbd9e194aef6