Pith. sign in

REVIEW 14 cited by

Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08861 v5 pith:FJGMMMB3 submitted 2024-09-13 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords fine-tuningmodelsrewardmatchingadjointcontroldiffusionexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning. In this work, we cast reward fine-tuning as stochastic optimal control (SOC). Critically, we prove that a very specific memoryless noise schedule must be enforced during fine-tuning, in order to account for the dependency between the noise variable and the generated samples. We also propose a new algorithm named Adjoint Matching which outperforms existing SOC algorithms, by casting SOC problems as a regression problem. We find that our approach significantly improves over existing methods for reward fine-tuning, achieving better consistency, realism, and generalization to unseen human preference reward models, while retaining sample diversity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Solving Inverse Problems with Flow-based Models via Model Predictive Control

    eess.IV 2026-01 conditional novelty 6.0 of 10

    MPC-Flow applies model predictive control to guide pretrained flow models through inverse problems, with a single-step variant that avoids backpropagation and scales to 32B-parameter models on consumer hardware.

  2. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Direct-Align and SRPO fine-tune FLUX using ground-truth-noise recovery and text-conditional relative rewards, improving human-evaluated realism and aesthetics roughly 3x.

  3. Feynman-Kac-Flow: Inference Steering of Conditional Flow Matching to an Energy-Tilted Posterior

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Feynman-Kac particle steering, previously diffusion-only, is derived for conditional flow matching and used to generate chirality-correct chemical transition states.

  4. FlowBack-Adjoint: Physics-Aware and Energy-Guided Conditional Flow-Matching for All-Atom Protein Backmapping

    physics.chem-ph 2025-08 conditional novelty 6.0 of 10

    FlowBack-Adjoint fine-tunes a flow-matching backmapping model with molecular mechanics energy gradients, reducing clashes and bond errors and producing lower-energy all-atom protein reconstructions.

  5. LSSGen: Leveraging Latent Space Scaling in Flow and Diffusion for Efficient Text to Image Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A latent-space scaling framework that replaces pixel-space upscaling with a trainable latent upsampler and noise compensation, yielding faster high-resolution text-to-image generation.

  6. Reinforcement Learning for Flow-Matching Policies

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Reward-weighted flow matching and GRPO with a learned reward surrogate both improve flow-matching policies beyond a suboptimal demonstrator on simulated unicycle tasks.

  7. Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Diffusion Tree Sampling is a Monte Carlo tree search over denoising trajectories that propagates terminal rewards backward to sample from reward-aligned distributions, showing up to 10x compute savings on tested benchmarks.

  8. Test-Time Scaling of Diffusion Models via Noise Trajectory Search

    cs.LG 2025-05 conditional novelty 6.0 of 10

    An epsilon-greedy search over per-step noise trajectories improves proxy rewards in diffusion image generation without retraining.

  9. Scaling Image and Video Generation via Test-Time Evolutionary Search

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Evolutionary search over denoising trajectories improves image and video generation quality and diversity as test-time compute increases, without retraining the generative model.

  10. Integration Matters: Rollout-Based Training for Constrained Diffusion Models

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Rollout-based fine-tuning with a learned adaptive guidance scaling yields near-zero constraint violations while preserving sample fidelity in constrained diffusion models.

  11. Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models

    cs.CV 2025-07 reject novelty 5.0 of 10

    Inversion-DPO uses DDIM inversion to convert winning and losing images into noise trajectories, yielding a simpler DPO loss for diffusion model alignment that trains faster and improves text-to-image and compositional...

  12. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...

  13. Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis

    cs.CV 2025-09 reject novelty 4.0 of 10

    A theoretical and empirical comparison claiming diffusion bridges have lower stochastic-optimal-control cost and greater robustness than flow matching when training data are scarce.

  14. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools