Pith. sign in

REVIEW 9 cited by

Robust Preference Optimization through Reward Model Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19316 v2 pith:4RWFQYLN submitted 2024-05-29 cs.LG cs.CL

classification cs.LGcs.CL
keywords preferencerewardmodeldistributionalignmentannotationsdatadistillation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language model (LM) post-training (or alignment) involves maximizing a reward function that is derived from preference annotations. Direct Preference Optimization (DPO) is a popular offline alignment method that trains a policy directly on preference data without the need to train a reward model or apply reinforcement learning. However, the empirical evidence suggests that DPO typically assigns implicit rewards that overfit, and trend towards infinite magnitude. This frequently leads to degenerate policies, sometimes causing even the probabilities of the preferred generations to go to zero. In this work, we analyze this phenomenon and use distillation to get a better proxy for the true preference distribution over generation pairs: we train the LM such that its induced implicit reward, i.e., the scaled log-likelihood ratio of the model to the reference model, matches an explicit reward model trained on the preference data. Moreover, to account for uncertainty in the reward model we are distilling from, we optimize against a family of reward models that, as a whole, is likely to include at least one reasonable proxy for the preference distribution. Our results show that distilling from such a family of reward models leads to improved robustness to distribution shift in preference annotations, while preserving the simple supervised nature of DPO.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

    eess.AS 2025-08 conditional novelty 6.0 of 10

    MPO improves TTS alignment by constructing multi-dimensional preference pairs and adding cross-entropy regularization to DPO, yielding better intelligibility, speaker similarity, and prosody.

  2. Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Applying DPO with preferences from people with intellectual disabilities improves German simplified-text readability but reduces semantic fidelity, and inconsistent target-group preferences prevent statistically signi...

  3. Learning a Pessimistic Reward Model in RLHF

    cs.LG 2025-05 reject novelty 6.0 of 10

    Pessimistic fine-tuning of reward models against rejection-sampling policies lets RLHF agents optimize greedily without KL regularization and still avoid reward hacking.

  4. Normalized Rewards for Preference Optimization

    cs.LG 2026-06 conditional novelty 5.0 of 10

    A regularization term that conserves the combined length-normalized probability of chosen and rejected responses reduces likelihood displacement in DPO/SimPO, improves AlpacaEval and benchmark outcomes, and acts prima...

  5. Risk-aware Direct Preference Optimization under Nested Risk Measure

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A token-level DPO variant that penalizes model drift with nested risk measures (CVaR and ERM) and reports improved alignment-drift tradeoffs.

  6. Frictional Agent Alignment Framework: Slow Down and Don't Break Things

    cs.CL 2025-05 reject novelty 5.0 of 10

    FAAF aligns an LLM with a two-part loss, one part conditioned on a detected belief-misalignment state, to generate reflection-prompting interventions, and reports higher win rates than DPO, IPO, and PPO on collaborati...

  7. Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    The submitted manuscript's abstract and full text are mismatched; the claimed 3D detection method is not present in the body.

  8. On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

    cs.LG 2025-05 reject novelty 4.0 of 10

    Symmetric losses preserve action rankings under symmetric label noise, and the paper's claim that they also handle asymmetric noise is invalid.

  9. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

Pith tools