REVIEW 7 cited by
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has emerged as a promising approach for alignment, acting as an RL-free alternative to Reinforcement Learning from Human Feedback (RLHF). Despite DPO's various advancements and inherent limitations, an in-depth review of these aspects is currently lacking in the literature. In this work, we present a comprehensive review of the challenges and opportunities in DPO, covering theoretical analyses, variants, relevant preference datasets, and applications. Specifically, we categorize recent studies on DPO based on key research questions to provide a thorough understanding of DPO's current landscape. Additionally, we propose several future research directions to offer insights on model alignment for the research community. An updated collection of relevant papers can be found on https://github.com/Mr-Loevan/DPO-Survey.
Forward citations
Cited by 7 Pith papers
-
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
TBPO posits a token-level Bradley-Terry model and derives a Bregman-divergence density-ratio matching loss that generalizes DPO while preserving token-level optimality.
-
Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs
ASR self-verification via best-of-N sampling eliminates observed catastrophic failures in multiple neural-codec TTS models, with distillation transferring most of the robustness to single-shot decoding.
-
Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization
ACE uses Direct Preference Optimization to learn sequential intervention policies, reporting 70–71% error reduction on a 5-node synthetic causal benchmark, with qualitative physics and economics demonstrations.
-
Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
OTPO uses unbalanced optimal transport to reweight token-level log-likelihood ratios in DPO, reporting up to 10.9% higher length-controlled win rate on AlpacaEval2 than DPO.
-
Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design
AAMFM combines ESM3, an antigen-geometry adapter, and Cal-DPO preference optimization rewarded by AlphaFold3-style scores to design antibody CDRs and structures, reporting higher predicted binding scores than prior methods.
-
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.
-
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy
AI copilot preference optimization is organized into a pre-, mid-, and post-interaction taxonomy, with a unified definition of AI copilots.
Discussion (0). Continue with ORCID to comment.