REVIEW 4 cited by
Boosting Latent Diffusion with Flow Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Visual synthesis has recently seen significant leaps in performance, largely due to breakthroughs in generative models. Diffusion models have been a key enabler, as they excel in image diversity. However, this comes at the cost of slow training and synthesis, which is only partially alleviated by latent diffusion. To this end, flow matching is an appealing approach due to its complementary characteristics of faster training and inference but less diverse synthesis. We demonstrate that introducing flow matching between a frozen diffusion model and a convolutional decoder enables high-resolution image synthesis at reduced computational cost and model size. A small diffusion model can then effectively provide the necessary visual diversity, while flow matching efficiently enhances resolution and detail by mapping the small to a high-dimensional latent space. These latents are then projected to high-resolution images by the subsequent convolutional decoder of the latent diffusion approach. Combining the diversity of diffusion models, the efficiency of flow matching, and the effectiveness of convolutional decoders, state-of-the-art high-resolution image synthesis is achieved at $1024^2$ pixels with minimal computational cost. Further scaling up our method we can reach resolutions up to $2048^2$ pixels. Importantly, our approach is orthogonal to recent approximation and speed-up strategies for the underlying model, making it easily integrable into the various diffusion model frameworks.
Forward citations
Cited by 4 Pith papers
-
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
CrossFlow turns text directly into images, and images into text, depth, and higher resolution, by flowing between modality latents without a noise prior or cross-attention.
-
DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow
DetailGen3D refines coarse 3D geometry into detailed geometry by learning a direct latent-space flow from coarse to fine shapes, guided by an input image.
-
Stable Flow: Vital Layers for Training-Free Image Editing
An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.
-
High-Resolution Image Synthesis via Next-Token Prediction
An autoregressive model with continuous tokens, a new positional embedding (VoPE), and a data-feedback training strategy achieves strong text-to-image benchmarks at resolutions up to 4K.
Discussion (0). Continue with ORCID to comment.