REVIEW 21 cited by
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-image diffusion models have recently emerged at the forefront of image generation, powered by very large-scale unsupervised or weakly supervised text-to-image training datasets. Due to their unsupervised training, controlling their behavior in downstream tasks, such as maximizing human-perceived image quality, image-text alignment, or ethical image generation, is difficult. Recent works finetune diffusion models to downstream reward functions using vanilla reinforcement learning, notorious for the high variance of the gradient estimators. In this paper, we propose AlignProp, a method that aligns diffusion models to downstream reward functions using end-to-end backpropagation of the reward gradient through the denoising process. While naive implementation of such backpropagation would require prohibitive memory resources for storing the partial derivatives of modern text-to-image models, AlignProp finetunes low-rank adapter weight modules and uses gradient checkpointing, to render its memory usage viable. We test AlignProp in finetuning diffusion models to various objectives, such as image-text semantic alignment, aesthetics, compressibility and controllability of the number of objects present, as well as their combinations. We show AlignProp achieves higher rewards in fewer training steps than alternatives, while being conceptually simpler, making it a straightforward choice for optimizing diffusion models for differentiable reward functions of interest. Code and Visualization results are available at https://align-prop.github.io/.
Forward citations
Cited by 21 Pith papers
-
AssetDropper: Asset Extraction via Diffusion Models with Reward-Driven Optimization
AssetDropper introduces a task-specific diffusion model, a 212k-pair synthetic dataset, and a generative reward model to extract standardized assets from reference images.
-
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
PURE achieves an Õ(√(d_R+d_F)/√N) suboptimality gap, up to horizon factors, in continuous-time RL with general function approximation, and adds low-switching and low-rollout variants.
-
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
SURE learns sample-adaptive variance in a latent reward model and uses that variance to weight dense post-training feedback, improving image and video diffusion alignment in reported experiments.
-
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
WorldCycle post-trains interactive video world models with reinforcement learning rewards for spatial closure and temporal consistency on reversible action cycles, reducing long-horizon drift and improving composite-a...
-
DanceOPD: On-Policy Generative Field Distillation
Hard-routed, single low-noise on-policy velocity matching composes conflicting image-generation capabilities into one flow student better than joint training, merging, or dense OPD baselines.
-
Plug-and-Play Guidance for Discrete Diffusion Models via Gradient-Informed Logit Correction
Introduces GILC, a training-free plug-and-play guidance framework for discrete diffusion models that uses Jacobian-free logit correction to achieve SOTA results on DNA, protein, and molecular generation tasks.
-
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
Diffusion image models can be aligned without human labels by supervising every denoising step with score targets from original versus degraded prompts.
-
Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
Optimizing the null-text embedding in classifier-free guidance aligns diffusion outputs to a target reward while preserving cross-reward quality.
-
A novel method and dataset for depth-guided image deblurring from smartphone Lidar
Lidar depth-guided deblurring via a zero-shot diffusion method, evaluated on a new 45-scene dataset, achieves the best perceptual quality (LPIPS).
-
ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning
ShortFT fine-tunes Stable Diffusion by backpropagating reward gradients through a distilled few-step shortcut denoising chain, improving alignment scores over DRaFT-LV and DRTune.
-
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
Diffusion Tree Sampling is a Monte Carlo tree search over denoising trajectories that propagates terminal rewards backward to sample from reward-aligned distributions, showing up to 10x compute savings on tested benchmarks.
-
Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
Decoupling visual similarity and attack success into separate diffusion-model alignment stages yields unrestricted adversarial images with state-of-the-art black-box transferability.
-
Local Manifold Approximation and Projection for Manifold-Aware Diffusion Planning
LoMAP projects each guided diffusion sample onto a PCA subspace of nearby offline trajectories, reducing infeasible plans and improving returns in Maze2D, MuJoCo locomotion, and AntMaze.
-
Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards
Text2Stereo adapts Stable Diffusion to generate wide-baseline stereo image pairs from text by fine-tuning with LoRA and a disparity-correlation consistency reward.
-
DiffusionReward: Enhancing Blind Face Restoration through Reward Feedback Learning
A reward-feedback fine-tuning framework trains a face reward model and uses its gradient plus structural and regularization losses to improve diffusion face restoration models.
-
Dichotomous Diffusion Policy Optimization
DIPOLE decomposes a KL-regularized RL objective into a pair of sigmoid-weighted diffusion policies whose score combination (CFG-like) yields stable and controllable policy improvement.
-
Instant Preference Alignment for Text-to-Image Diffusion Models
An MLLM-driven, training-free pipeline extracts preference keywords from a reference image and modulates diffusion cross-attention at global and regional levels for instant, multi-round preference-aligned image generation.
-
ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models
ImageReFL combines base-model early diffusion steps with a real-image-based fine-tuning objective to improve the quality-diversity trade-off in reward-aligned text-to-image generation.
-
$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion
A pairwise-conditioned diffusion model generates instructional illustrations from procedural text and is finetuned with a text-image alignment reward.
-
Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences
SmPO-Diffusion improves diffusion-model preference alignment with reward-model soft labels and ReNoise inversion, reporting higher human-preference scores and up to 26x lower training cost than Diffusion-KTO.
-
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.
Discussion (0). Continue with ORCID to comment.