Pith. sign in

REVIEW 7 cited by

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.15595 v4 pith:7E5YFW22 submitted 2024-10-21 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords preferenceresearchalignmentapplicationscomprehensivedatasetsdirecthuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has emerged as a promising approach for alignment, acting as an RL-free alternative to Reinforcement Learning from Human Feedback (RLHF). Despite DPO's various advancements and inherent limitations, an in-depth review of these aspects is currently lacking in the literature. In this work, we present a comprehensive review of the challenges and opportunities in DPO, covering theoretical analyses, variants, relevant preference datasets, and applications. Specifically, we categorize recent studies on DPO based on key research questions to provide a thorough understanding of DPO's current landscape. Additionally, we propose several future research directions to offer insights on model alignment for the research community. An updated collection of relevant papers can be found on https://github.com/Mr-Loevan/DPO-Survey.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    TBPO posits a token-level Bradley-Terry model and derives a Bregman-divergence density-ratio matching loss that generalizes DPO while preserving token-level optimality.

  2. Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    ASR self-verification via best-of-N sampling eliminates observed catastrophic failures in multiple neural-codec TTS models, with distillation transferring most of the robustness to single-shot decoding.

  3. Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization

    cs.LG 2026-02 conditional novelty 6.0 of 10

    ACE uses Direct Preference Optimization to learn sequential intervention policies, reporting 70–71% error reduction on a 5-node synthetic causal benchmark, with qualitative physics and economics demonstrations.

  4. Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    OTPO uses unbalanced optimal transport to reweight token-level log-likelihood ratios in DPO, reporting up to 10.9% higher length-controlled win rate on AlpacaEval2 than DPO.

  5. Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design

    q-bio.BM 2026-07 reject novelty 5.0 of 10

    AAMFM combines ESM3, an antigen-geometry adapter, and Cal-DPO preference optimization rewarded by AlphaFold3-style scores to design antibody CDRs and structures, reporting higher predicted binding scores than prior methods.

  6. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.

  7. Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy

    cs.AI 2025-05 reject novelty 4.0 of 10

    AI copilot preference optimization is organized into a pre-, mid-, and post-interaction taxonomy, with a unified definition of AI copilots.

Pith tools