Pith. sign in

REVIEW 6 cited by

Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.00847 v2 pith:VLFXKXM4 submitted 2024-10-01 cs.LG

classification cs.LG
keywords rewardmodelsdemonstratehumanuncertaintyurmecaptureensemble
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reward models (RMs) are essential for aligning large language models (LLM) with human expectations. However, existing RMs struggle to capture the stochastic and uncertain nature of human preferences and fail to assess the reliability of reward predictions. To address these challenges, we introduce the Uncertainty-aware Reward Model (URM) and its ensemble variant, URME. URM employs a probabilistic value head to capture aleatoric uncertainty by modeling the distribution of disentangled human preference attributes. URME further quantifies epistemic uncertainty by examining discrepancies among individual URMs within the ensemble, enabling identification of unreliable evaluations. Our empirical evaluations demonstrate that URM achieves strong performance on RewardBench, outperforming competitive large-scale models. Additionally, extensive experiments, including best-of-n sampling (BoN), iterative direct preference optimization (iterative DPO), and proximal policy optimization (PPO), demonstrate that URM and URME significantly enhance LLMs' generation quality. Notably, reward predictions with lower uncertainty are far more reliable, demonstrate significantly higher quality, and result in substantially improved alignment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Fusing hidden states of multiple frozen reward models under shared-private constraints improves pairwise preference accuracy over scalar-score fusion and single-expert adapters on LM-Arena and PPE.

  2. Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

    cs.CL 2025-07 conditional novelty 6.0 of 10

    RLSF uses a model's own answer-span confidence as an intrinsic reward to create preference data, then applies DPO or PPO to improve calibration and reasoning without external labels.

  3. Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Dialog-act and maxim-aware prompting improves LLM judge accuracy on multi-turn preference data by up to 8 points, with further gains from jury-style voting.

  4. Decision Flow Policy Optimization

    cs.LG 2025-05 reject novelty 6.0 of 10

    Decision Flow frames the gradual action generation of flow-based policies as a flow MDP and updates the flow policy with flow-level value functions, reporting state-of-the-art results on several D4RL tasks.

  5. Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Repeated sampling with a verifier improves open-ended multilingual generation, and reward-based verifiers are needed for math and code tasks.

  6. Towards Reliable, Uncertainty-Aware Alignment

    cs.LG 2025-07 reject novelty 4.0 of 10

    Variance-aware RLHF adds a variance-weighted KL penalty to PPO and reduces reward variance and the risk of underperforming the reference policy in the paper's experiments.

Pith tools