REVIEW 14 cited by
Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning. In this work, we cast reward fine-tuning as stochastic optimal control (SOC). Critically, we prove that a very specific memoryless noise schedule must be enforced during fine-tuning, in order to account for the dependency between the noise variable and the generated samples. We also propose a new algorithm named Adjoint Matching which outperforms existing SOC algorithms, by casting SOC problems as a regression problem. We find that our approach significantly improves over existing methods for reward fine-tuning, achieving better consistency, realism, and generalization to unseen human preference reward models, while retaining sample diversity.
Forward citations
Cited by 14 Pith papers
-
Solving Inverse Problems with Flow-based Models via Model Predictive Control
MPC-Flow applies model predictive control to guide pretrained flow models through inverse problems, with a single-step variant that avoids backpropagation and scales to 32B-parameter models on consumer hardware.
-
Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
Direct-Align and SRPO fine-tune FLUX using ground-truth-noise recovery and text-conditional relative rewards, improving human-evaluated realism and aesthetics roughly 3x.
-
Feynman-Kac-Flow: Inference Steering of Conditional Flow Matching to an Energy-Tilted Posterior
Feynman-Kac particle steering, previously diffusion-only, is derived for conditional flow matching and used to generate chirality-correct chemical transition states.
-
FlowBack-Adjoint: Physics-Aware and Energy-Guided Conditional Flow-Matching for All-Atom Protein Backmapping
FlowBack-Adjoint fine-tunes a flow-matching backmapping model with molecular mechanics energy gradients, reducing clashes and bond errors and producing lower-energy all-atom protein reconstructions.
-
LSSGen: Leveraging Latent Space Scaling in Flow and Diffusion for Efficient Text to Image Generation
A latent-space scaling framework that replaces pixel-space upscaling with a trainable latent upsampler and noise compensation, yielding faster high-resolution text-to-image generation.
-
Reinforcement Learning for Flow-Matching Policies
Reward-weighted flow matching and GRPO with a learned reward surrogate both improve flow-matching policies beyond a suboptimal demonstrator on simulated unicycle tasks.
-
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
Diffusion Tree Sampling is a Monte Carlo tree search over denoising trajectories that propagates terminal rewards backward to sample from reward-aligned distributions, showing up to 10x compute savings on tested benchmarks.
-
Test-Time Scaling of Diffusion Models via Noise Trajectory Search
An epsilon-greedy search over per-step noise trajectories improves proxy rewards in diffusion image generation without retraining.
-
Scaling Image and Video Generation via Test-Time Evolutionary Search
Evolutionary search over denoising trajectories improves image and video generation quality and diversity as test-time compute increases, without retraining the generative model.
-
Integration Matters: Rollout-Based Training for Constrained Diffusion Models
Rollout-based fine-tuning with a learned adaptive guidance scaling yields near-zero constraint violations while preserving sample fidelity in constrained diffusion models.
-
Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
Inversion-DPO uses DDIM inversion to convert winning and losing images into noise trajectories, yielding a simpler DPO loss for diffusion model alignment that trains faster and improves text-to-image and compositional...
-
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...
-
Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
A theoretical and empirical comparison claiming diffusion bridges have lower stochastic-optimal-control cost and greater robustness than flow matching when training data are scarce.
-
Reinforcement Learning: From Algorithms To Foundation Models
A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.
Discussion (0). Continue with ORCID to comment.