Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2410.14872.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.358755Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T06:25:27.821937Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 1be5e7df-2b31-4e5b-af58-5138884f6000 · inbound
Qwen2.5 Technical Report How to Evaluate Reward Models for RLHF
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 245b1fac-f3ff-4cb2-bdb8-45ead85e4fab · inbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators How to Evaluate Reward Models for RLHF
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9676dbaa-35e9-43c3-98e6-a78b00c7fb3e · inbound
On the Robustness of Reward Models for Language Model Alignment How to Evaluate Reward Models for RLHF
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246eac89-1703-458f-8d1c-c5afa4793d67 · inbound
WorldPM: Scaling Human Preference Modeling How to Evaluate Reward Models for RLHF
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32242d9b-6592-4ddc-8418-e2dfbe177a3e · inbound
RewardBench 2: Advancing Reward Model Evaluation How to Evaluate Reward Models for RLHF
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 284e87a8-4613-42db-9312-a2f323338b5d · inbound
RewardAnything: Generalizable Principle-Following Reward Models How to Evaluate Reward Models for RLHF
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dafde6d9-7c3b-4dae-ac77-f8f20c1952cf · inbound
Bridging Offline and Online Reinforcement Learning for LLMs How to Evaluate Reward Models for RLHF
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c68ee6-2f9c-4abb-964f-6da8ef9c2e91 · inbound
Active Query Selection for Crowd-Based Reinforcement Learning How to Evaluate Reward Models for RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03dbc1b8-32d0-41cb-8622-38e8d95efb1e · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data How to Evaluate Reward Models for RLHF
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 195b989d-6a45-47a2-9d0e-40a5a1bab2c3 · inbound
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems How to Evaluate Reward Models for RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.