REVIEW 20 cited by
Diffusion Model Alignment Using Direct Preference Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs, human preference learning has not been widely explored in text-to-image diffusion models; the best existing approach is to fine-tune a pretrained model using carefully curated high quality images and captions to improve visual appeal and text alignment. We propose Diffusion-DPO, a method to align diffusion models to human preferences by directly optimizing on human comparison data. Diffusion-DPO is adapted from the recently developed Direct Preference Optimization (DPO), a simpler alternative to RLHF which directly optimizes a policy that best satisfies human preferences under a classification objective. We re-formulate DPO to account for a diffusion model notion of likelihood, utilizing the evidence lower bound to derive a differentiable objective. Using the Pick-a-Pic dataset of 851K crowdsourced pairwise preferences, we fine-tune the base model of the state-of-the-art Stable Diffusion XL (SDXL)-1.0 model with Diffusion-DPO. Our fine-tuned base model significantly outperforms both base SDXL-1.0 and the larger SDXL-1.0 model consisting of an additional refinement model in human evaluation, improving visual appeal and prompt alignment. We also develop a variant that uses AI feedback and has comparable performance to training on human preferences, opening the door for scaling of diffusion model alignment methods.
Forward citations
Cited by 20 Pith papers
-
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.
-
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.
-
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
SURE learns sample-adaptive variance in a latent reward model and uses that variance to weight dense post-training feedback, improving image and video diffusion alignment in reported experiments.
-
Learning Sampling Parameters for Diffusion Models
An LLM policy trained with GRPO can emit prompt-conditioned, timestep-varying diffusion sampling parameters that beat fixed defaults and prior LLM schedulers on preference metrics.
-
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
CACFM applies RL to adaptively select critical regions in probability flow ODE trajectories for consistency distillation, yielding SOTA few-step results on FLUX and SDXL.
-
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
Phys4D improves video diffusion models' physical consistency via pseudo-supervised depth/motion heads, simulation-grounded fine-tuning, and RL with a 4D Chamfer reward.
-
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.
-
Test-Time Scaling of Diffusion Models via Noise Trajectory Search
An epsilon-greedy search over per-step noise trajectories improves proxy rewards in diffusion image generation without retraining.
-
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.
-
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.
-
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Two plug-and-play strategies — per-timestep advantage weighting and advantage-based trajectory replay — improve diffusion RLHF sample efficiency up to 6× across five reward functions.
-
MolFORM: Multi-modal Flow Matching for Structure-Based Drug Design
A flow-matching model with direct preference optimization fine-tuning generates protein-binding molecules faster than diffusion baselines, with improved docking scores on the CrossDocked2020 benchmark.
-
ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models
ImageReFL combines base-model early diffusion steps with a real-image-based fine-tuning objective to improve the quality-diversity trade-off in reward-aligned text-to-image generation.
-
$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion
A pairwise-conditioned diffusion model generates instructional illustrations from procedural text and is finetuned with a text-image alignment reward.
-
Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation
MuDi-Pro fine-tunes a multi-guided diffusion transformer with DPO using guidance-score preferences to improve controllability of traffic scenario generation on nuScenes.
-
YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment
Introduces a six-axis contradictory-objective benchmark and a weighted DPO variant (CAO), but the claimed balanced alignment rests on evaluations that reuse the training objectives.
-
CoDe: Blockwise Control for Denoising Diffusion Models
CoDe applies blockwise best-of-N sampling during diffusion denoising, with Tweedie-based reward estimates, to align generated images to differentiable or non-differentiable rewards.
-
Constraints-Guided Diffusion Reasoner for Neuro-Symbolic Learning
A masked diffusion model fine-tuned with group-reward reinforcement learning solves Sudoku and Maze puzzles more accurately than the same model trained by supervised learning alone.
-
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.
-
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
A DPO variant that kernelizes the preference loss and swaps KL for other divergences is claimed to improve alignment, but the math and evaluation do not support the state-of-the-art claim.
Discussion (0). Continue with ORCID to comment.