Pith. sign in

REVIEW 7 cited by

How to Evaluate Reward Models for RLHF

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14872 v2 pith:3M75WMYO submitted 2024-10-18 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords rewardmodelpreferencerlhfhumanperformancedownstreammodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-standard approach is to run a full RLHF training pipeline and directly probe downstream LLM performance. However, this process is prohibitively expensive. To address this, we build a predictive model of downstream LLM performance by evaluating the reward model on proxy tasks. These proxy tasks consist of a large-scale human preference and a verifiable correctness preference dataset, in which we measure 12 metrics across 12 domains. To investigate which reward model metrics are most correlated to gold-standard RLHF outcomes, we launch an end-to-end RLHF experiment on a large-scale crowdsourced human preference platform to view real reward model downstream performance as ground truth. Ultimately, we compile our data and findings into Preference Proxy Evaluations (PPE), the first reward model benchmark explicitly linked to post-RLHF real-world human preference performance, which we open-source for public use and further development. Our code and evaluations can be found at https://github.com/lmarena/PPE .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

    cs.CL 2025-04 conditional novelty 7.0 of 10

    JETTS, a new benchmark, shows LLM-as-judges are competitive in response reranking, worse than process reward models in beam search, and ineffective as critique providers for refinement.

  2. Bridging Offline and Online Reinforcement Learning for LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Semi-online DPO, which syncs the generation model every few update steps, performs nearly as well as fully online DPO and GRPO, while strongly beating offline DPO.

  3. RewardAnything: Generalizable Principle-Following Reward Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RewardAnything follows natural-language reward principles at inference time and, with the new RABench benchmark, demonstrates that principle-conditioned listwise training beats fixed-preference reward models on held-o...

  4. WorldPM: Scaling Human Preference Modeling

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Preference modeling exhibits scaling laws for objective and adversarial tasks, with up to 5-14% gains when used as initialization for fine-tuning on smaller human-preference datasets.

  5. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  6. On the Robustness of Reward Models for Language Model Alignment

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Adding a batch-wise sum-to-zero penalty to Bradley-Terry reward modeling makes reward models more robust to unseen prompts and responses, according to experiments across multiple model families and benchmarks.

  7. Active Query Selection for Crowd-Based Reinforcement Learning

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Extending the Advise algorithm with variational crowd modelling and entropy-based query selection yields faster learning in small tabular RL tasks, especially highly constrained ones.

Pith tools