REVIEW 8 cited by
Text-to-Image Rectified Flow as Plug-and-Play Priors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale diffusion models have achieved remarkable performance in generative tasks. Beyond their initial training applications, these models have proven their ability to function as versatile plug-and-play priors. For instance, 2D diffusion models can serve as loss functions to optimize 3D implicit models. Rectified flow, a novel class of generative models, enforces a linear progression from the source to the target distribution and has demonstrated superior performance across various domains. Compared to diffusion-based methods, rectified flow approaches surpass in terms of generation quality and efficiency, requiring fewer inference steps. In this work, we present theoretical and experimental evidence demonstrating that rectified flow based methods offer similar functionalities to diffusion models - they can also serve as effective priors. Besides the generative capabilities of diffusion priors, motivated by the unique time-symmetry properties of rectified flow models, a variant of our method can additionally perform image inversion. Experimentally, our rectified flow-based priors outperform their diffusion counterparts - the SDS and VSD losses - in text-to-3D generation. Our method also displays competitive performance in image inversion and editing.
Forward citations
Cited by 8 Pith papers
-
Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
Functional rectified flow is defined and proved to preserve marginals in separable Hilbert spaces, with functional flow matching and probability-flow ODEs as special cases.
-
Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
RoMaP enables precise and drastic part-level edits in 3D Gaussian scenes using SH-based soft-label 3D segmentation and a regularized SDS loss anchored on scheduled latent-mixing images.
-
Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection
Z-Sampling alternates high-guidance denoising and low-guidance inversion at each step to improve prompt alignment in pretrained text-to-image diffusion models.
-
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
FlowChef enables training-free, inversion-free, backprop-free controlled generation for rectified flow models by replacing the gradient through the model with the direct loss gradient on the estimated clean image.
-
DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh
DAGSM is a text-to-3D avatar pipeline that generates body and garments as separate mesh-bound 2DGS models, enabling clothing replacement, texture editing, and animatable cloth.
-
Translationese as a Rational Response to Translation Task Difficulty
Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.
-
FlowSteer: Conditioning Flow Field for Consistent Image Restoration
A sparse mid-to-late schedule of null-space fidelity updates lets a frozen text-to-image flow model restore images with high measurement consistency.
-
FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.
Discussion (0). Continue with ORCID to comment.