Pith. sign in

REVIEW 12 cited by

Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04314 v3 pith:PNTSPNDI submitted 2024-06-06 cs.CV

classification cs.CV
keywords preferenceaestheticsdiffusionmodelsaestheticlabelsmethodsdenoising
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating visually appealing images is fundamental to modern text-to-image generation models. A potential solution to better aesthetics is direct preference optimization (DPO), which has been applied to diffusion models to improve general image quality including prompt alignment and aesthetics. Popular DPO methods propagate preference labels from clean image pairs to all the intermediate steps along the two generation trajectories. However, preference labels provided in existing datasets are blended with layout and aesthetic opinions, which would disagree with aesthetic preference. Even if aesthetic labels were provided (at substantial cost), it would be hard for the two-trajectory methods to capture nuanced visual differences at different steps. To improve aesthetics economically, this paper uses existing generic preference data and introduces step-by-step preference optimization (SPO) that discards the propagation strategy and allows fine-grained image details to be assessed. Specifically, at each denoising step, we 1) sample a pool of candidates by denoising from a shared noise latent, 2) use a step-aware preference model to find a suitable win-lose pair to supervise the diffusion model, and 3) randomly select one from the pool to initialize the next denoising step. This strategy ensures that diffusion models focus on the subtle, fine-grained visual differences instead of layout aspect. We find that aesthetics can be significantly enhanced by accumulating these improved minor differences. When fine-tuning Stable Diffusion v1.5 and SDXL, SPO yields significant improvements in aesthetics compared with existing DPO methods while not sacrificing image-text alignment compared with vanilla models. Moreover, SPO converges much faster than DPO methods due to the use of more correct preference labels provided by the step-aware preference model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.

  2. ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ShortFT fine-tunes Stable Diffusion by backpropagating reward gradients through a distilled few-step shortcut denoising chain, improving alignment scores over DRaFT-LV and DRTune.

  3. X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GRPO reinforcement learning applied to a discrete autoregressive image generator with a diffusion decoder improves instruction following, image quality, and long-text rendering in a unified multimodal model.

  4. Fake it till You Make it: Reward Modeling as Discriminative Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GAN-RM trains a CLIP-based discriminator to distinguish a few hundred preference proxy images from model outputs, then uses it for Best-of-N selection, SFT, and DPO.

  5. Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Decoupling visual similarity and attack success into separate diffusion-model alignment stages yields unrestricted adversarial images with state-of-the-art black-box transferability.

  6. Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences

    physics.ao-ph 2025-05 conditional novelty 6.0 of 10

    Align-DA uses direct preference optimization to align a score-based data assimilation prior with assimilation accuracy, forecast skill, and physical adherence rewards, improving analysis quality over unaligned diffusi...

  7. Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CoCA redistributes a single final image reward across denoising steps using cosine similarity between intermediate and final latents, improving RL fine-tuning sample efficiency on four human preference rewards.

  8. Scaling Image and Video Generation via Test-Time Evolutionary Search

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Evolutionary search over denoising trajectories improves image and video generation quality and diversity as test-time compute increases, without retraining the generative model.

  9. Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models

    cs.CV 2025-07 reject novelty 5.0 of 10

    Inversion-DPO uses DDIM inversion to convert winning and losing images into noise trajectories, yielding a simpler DPO loss for diffusion model alignment that trains faster and improves text-to-image and compositional...

  10. GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GigaVideo-1 fine-tunes Wan2.1 on synthetic weakness-targeted prompts with VLM reward reweighting and reports ~4% average VBench-2.0 gains per dimension at 4 GPU-hours each, though joint training gains less.

  11. ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ImageReFL combines base-model early diffusion steps with a real-image-based fine-tuning objective to improve the quality-diversity trade-off in reward-aligned text-to-image generation.

  12. Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Blending a base diffusion model with its RL-finetuned version at sampling time lets users dial alignment strength, with the blend weight corresponding to the KL-regularization coefficient beta/w.

Pith tools