Pith. sign in

REVIEW 4 cited by

Boosting Latent Diffusion with Flow Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07360 v3 pith:25KTQ7JY submitted 2023-12-12 cs.CV

classification cs.CV
keywords diffusionflowmatchingmodelsynthesislatentapproachconvolutional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Visual synthesis has recently seen significant leaps in performance, largely due to breakthroughs in generative models. Diffusion models have been a key enabler, as they excel in image diversity. However, this comes at the cost of slow training and synthesis, which is only partially alleviated by latent diffusion. To this end, flow matching is an appealing approach due to its complementary characteristics of faster training and inference but less diverse synthesis. We demonstrate that introducing flow matching between a frozen diffusion model and a convolutional decoder enables high-resolution image synthesis at reduced computational cost and model size. A small diffusion model can then effectively provide the necessary visual diversity, while flow matching efficiently enhances resolution and detail by mapping the small to a high-dimensional latent space. These latents are then projected to high-resolution images by the subsequent convolutional decoder of the latent diffusion approach. Combining the diversity of diffusion models, the efficiency of flow matching, and the effectiveness of convolutional decoders, state-of-the-art high-resolution image synthesis is achieved at $1024^2$ pixels with minimal computational cost. Further scaling up our method we can reach resolutions up to $2048^2$ pixels. Importantly, our approach is orthogonal to recent approximation and speed-up strategies for the underlying model, making it easily integrable into the various diffusion model frameworks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

    cs.CV 2024-12 conditional novelty 7.0 of 10

    CrossFlow turns text directly into images, and images into text, depth, and higher resolution, by flowing between modality latents without a noise prior or cross-attention.

  2. DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow

    cs.CV 2024-11 conditional novelty 6.0 of 10

    DetailGen3D refines coarse 3D geometry into detailed geometry by learning a direct latent-space flow from coarse to fine shapes, guided by an input image.

  3. Stable Flow: Vital Layers for Training-Free Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.

  4. High-Resolution Image Synthesis via Next-Token Prediction

    cs.CV 2024-11 conditional novelty 4.0 of 10

    An autoregressive model with continuous tokens, a new positional embedding (VoPE), and a data-feedback training strategy achieves strong text-to-image benchmarks at resolutions up to 4K.

Pith tools