Pith. sign in

REVIEW 20 cited by

Diffusion Model Alignment Using Direct Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12908 v1 pith:QZNDROGD submitted 2023-11-21 cs.CV cs.AIcs.GRcs.LG

classification cs.CVcs.AIcs.GRcs.LG
keywords modelhumandiffusionpreferencesalignmentbasediffusion-dpomodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs, human preference learning has not been widely explored in text-to-image diffusion models; the best existing approach is to fine-tune a pretrained model using carefully curated high quality images and captions to improve visual appeal and text alignment. We propose Diffusion-DPO, a method to align diffusion models to human preferences by directly optimizing on human comparison data. Diffusion-DPO is adapted from the recently developed Direct Preference Optimization (DPO), a simpler alternative to RLHF which directly optimizes a policy that best satisfies human preferences under a classification objective. We re-formulate DPO to account for a diffusion model notion of likelihood, utilizing the evidence lower bound to derive a differentiable objective. Using the Pick-a-Pic dataset of 851K crowdsourced pairwise preferences, we fine-tune the base model of the state-of-the-art Stable Diffusion XL (SDXL)-1.0 model with Diffusion-DPO. Our fine-tuned base model significantly outperforms both base SDXL-1.0 and the larger SDXL-1.0 model consisting of an additional refinement model in human evaluation, improving visual appeal and prompt alignment. We also develop a variant that uses AI feedback and has comparable performance to training on human preferences, opening the door for scaling of diffusion model alignment methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.

  2. Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

    cs.CV 2025-05 conditional novelty 7.0 of 10

    PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.

  3. Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

    cs.CV 2026-08 conditional novelty 6.0 of 10

    SURE learns sample-adaptive variance in a latent reward model and uses that variance to weight dense post-training feedback, improving image and video diffusion alignment in reported experiments.

  4. Learning Sampling Parameters for Diffusion Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An LLM policy trained with GRPO can emit prompt-conditioned, timestep-varying diffusion sampling parameters that beat fixed defaults and prior LLM schedulers on preference metrics.

  5. Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CACFM applies RL to adaptively select critical regions in probability flow ODE trajectories for consistency distillation, yielding SOTA few-step results on FLUX and SDXL.

  6. Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Phys4D improves video diffusion models' physical consistency via pseudo-supervised depth/motion heads, simulation-grounded fine-tuning, and RL with a 4D Chamfer reward.

  7. JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

    cs.SD 2025-07 conditional novelty 6.0 of 10

    JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.

  8. Test-Time Scaling of Diffusion Models via Noise Trajectory Search

    cs.LG 2025-05 conditional novelty 6.0 of 10

    An epsilon-greedy search over per-step noise trajectories improves proxy rewards in diffusion image generation without retraining.

  9. TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.

  10. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.

  11. Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Two plug-and-play strategies — per-timestep advantage weighting and advantage-based trajectory replay — improve diffusion RLHF sample efficiency up to 6× across five reward functions.

  12. MolFORM: Multi-modal Flow Matching for Structure-Based Drug Design

    cs.CE 2025-07 conditional novelty 5.0 of 10

    A flow-matching model with direct preference optimization fine-tuning generates protein-binding molecules faster than diffusion baselines, with improved docking scores on the CrossDocked2020 benchmark.

  13. ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ImageReFL combines base-model early diffusion steps with a real-image-based fine-tuning objective to improve the quality-diversity trade-off in reward-aligned text-to-image generation.

  14. $I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

    cs.CL 2025-05 reject novelty 5.0 of 10

    A pairwise-conditioned diffusion model generates instructional illustrations from procedural text and is finetuned with a text-image alignment reward.

  15. Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation

    cs.LG 2025-02 reject novelty 5.0 of 10

    MuDi-Pro fine-tunes a multi-guided diffusion transformer with DPO using guidance-score preferences to improve controllability of traffic scenario generation on nuScenes.

  16. YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing Multi-Objective Optimization based DPO for Text-to-Image Alignment

    cs.AI 2025-02 reject novelty 5.0 of 10

    Introduces a six-axis contradictory-objective benchmark and a weighted DPO variant (CAO), but the claimed balanced alignment rests on evaluations that reuse the training objectives.

  17. CoDe: Blockwise Control for Denoising Diffusion Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CoDe applies blockwise best-of-N sampling during diffusion denoising, with Tweedie-based reward estimates, to align generated images to differentiable or non-differentiable rewards.

  18. Constraints-Guided Diffusion Reasoner for Neuro-Symbolic Learning

    cs.AI 2025-08 conditional novelty 4.0 of 10

    A masked diffusion model fine-tuned with group-reward reinforcement learning solves Sudoku and Maze puzzles more accurately than the same model trained by supervised learning alone.

  19. Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

    cs.CV 2025-05 reject novelty 4.0 of 10

    Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.

  20. DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

    cs.LG 2025-01 reject novelty 4.0 of 10

    A DPO variant that kernelizes the preference loss and swaps KL for other divergences is claimed to improve alignment, but the math and evaluation do not support the state-of-the-art claim.

Pith tools