Pith. sign in

REVIEW 3 cited by

Filtered Direct Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13846 v4 pith:5BYBUN4A submitted 2024-04-22 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords preferencequalitydatasetrlhfdirectfdpomodeloptimization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning from human feedback (RLHF) plays a crucial role in aligning language models with human preferences. While the significance of dataset quality is generally recognized, explicit investigations into its impact within the RLHF framework, to our knowledge, have been limited. This paper addresses the issue of text quality within the preference dataset by focusing on direct preference optimization (DPO), an increasingly adopted reward-model-free RLHF method. We confirm that text quality significantly influences the performance of models optimized with DPO more than those optimized with reward-model-based RLHF. Building on this new insight, we propose an extension of DPO, termed filtered direct preference optimization (fDPO). fDPO uses a trained reward model to monitor the quality of texts within the preference dataset during DPO training. Samples of lower quality are discarded based on comparisons with texts generated by the model being optimized, resulting in a more accurate dataset. Experimental results demonstrate that fDPO enhances the final model performance. Our code is available at https://github.com/CyberAgentAILab/filtered-dpo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A preference optimization strategy using confidence-based hard example mining, similarity retrieval, and synthetic counterfactual rationales improves chest X-ray VQA accuracy by 8.93% relative over supervised fine-tuning.

  2. From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations

    cs.CL 2025-07 unverdicted novelty 6.0 of 10

    DeFactoX, a curriculum-driven DPO variant with Actuality and Finesse loss weighting, improves automatic and human scores for Hindi news explanation generation over existing preference optimization baselines.

  3. Influence Functions for Preference Dataset Pruning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Conjugate-gradient influence functions can mildly improve reward-model accuracy after pruning 10% of a preference dataset, but the gain is not statistically significant and gradient similarity better identifies helpfu...

Pith tools