REVIEW 19 cited by
Flow Matching in Latent Space
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Flow matching is a recent framework to train generative models that exhibits impressive empirical performance while being relatively easier to train compared with diffusion-based models. Despite its advantageous properties, prior methods still face the challenges of expensive computing and a large number of function evaluations of off-the-shelf solvers in the pixel space. Furthermore, although latent-based generative methods have shown great success in recent years, this particular model type remains underexplored in this area. In this work, we propose to apply flow matching in the latent spaces of pretrained autoencoders, which offers improved computational efficiency and scalability for high-resolution image synthesis. This enables flow-matching training on constrained computational resources while maintaining their quality and flexibility. Additionally, our work stands as a pioneering contribution in the integration of various conditions into flow matching for conditional generation tasks, including label-conditioned image generation, image inpainting, and semantic-to-image generation. Through extensive experiments, our approach demonstrates its effectiveness in both quantitative and qualitative results on various datasets, such as CelebA-HQ, FFHQ, LSUN Church & Bedroom, and ImageNet. We also provide a theoretical control of the Wasserstein-2 distance between the reconstructed latent flow distribution and true data distribution, showing it is upper-bounded by the latent flow matching objective. Our code will be available at https://github.com/VinAIResearch/LFM.git.
Forward citations
Cited by 19 Pith papers
-
Latent Flow Matching for Arbitrage-Aware Implied Volatility Surface Generation
Latent flow matching with an arbitrage-regularized VAE generates implied volatility surfaces that match the empirical distribution and pass static no-arbitrage tests at a higher rate than GAN and diffusion baselines.
-
Dense Temporal Contrast Synthesis via Conditioned Latent Transport
A single-pass latent transport model generates dense DCE-MRI contrast time series from pre-contrast scans, improving downstream tumor segmentation (Dice 0.60 vs 0.49) and preserving management decisions in 70% of read...
-
ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching
ZipL-Dialog cuts peak GPU memory 11.22× and speeds inference 2.23× for multi-minute zero-shot dialog TTS by doing conditional flow matching in a 4× compressed latent space while keeping perceptual naturalness.
-
DanceOPD: On-Policy Generative Field Distillation
Hard-routed, single low-noise on-policy velocity matching composes conflicting image-generation capabilities into one flow student better than joint training, merging, or dense OPD baselines.
-
LSSGen: Leveraging Latent Space Scaling in Flow and Diffusion for Efficient Text to Image Generation
A latent-space scaling framework that replaces pixel-space upscaling with a trainable latent upsampler and noise compensation, yielding faster high-resolution text-to-image generation.
-
Hierarchical Rectified Flow Matching with Mini-Batch Couplings
Mini-batch couplings in data and velocity space simplify the hierarchy of velocity distributions in hierarchical rectified flow matching, improving low-step generation quality.
-
Latent Thermodynamic Flows: Unified Representation Learning and Generative Modeling of Temperature-Dependent Behaviors from Limited Data
LaTF combines state-predictive information bottleneck with normalizing flows and a temperature-steerable tilted Gaussian prior to infer free energy surfaces at unseen temperatures from simulation data at two temperatures.
-
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
DisCon treats discrete image tokens as conditioning signals rather than targets, letting a continuous autoregressive model refine details and reach gFID 1.38 on ImageNet-256.
-
FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
FlowRAM pairs a shrinking 3D attention region with flow-matching action generation and a Mamba fusion model, setting new RLBench state-of-the-art results.
-
Decision Flow Policy Optimization
Decision Flow frames the gradual action generation of flow-based policies as a flow MDP and updates the flow policy with flow-level value functions, reporting state-of-the-art results on several D4RL tasks.
-
Designing a Conditional Prior Distribution for Flow-Based Generative Models
Condition-specific Gaussian mixture priors shorten flow-matching paths and improve FID, KID, and CLIP scores at low sampling steps on ImageNet-64 and MS-COCO.
-
Spectral Consistent Flow for One-step 3D Medical Image Translation
A one-step latent Brownian-bridge flow plus frequency-domain gain correction produces more accurate 3D medical image translations than multi-step diffusion and prior single-step baselines across four datasets.
-
Straight-Path Flow Matching for Incomplete Multi-View Clustering
Straight-path flow matching between paired latent representations outperforms diffusion-based methods for incomplete multi-view clustering by preserving cluster structure during view completion.
-
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
FaME uses stored trajectories of IQA-selected low-quality images as negative guidance to improve perceptual quality of diffusion generation while keeping FID.
-
CloudBreaker: Breaking the Cloud Covers of Sentinel-2 Images using Multi-Stage Trained Conditional Flow Matching on Sentinel-1
CloudBreaker trains a multi-stage conditional flow-matching model on paired Sentinel-1 radar and Sentinel-2 optical data to synthesize RGB, NDVI, and NDWI images under cloud cover.
-
Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models
Normalized Attention Guidance (NAG) stabilizes attention-space extrapolation with L1 normalization and refinement, restoring negative prompting in few-step diffusion models across architectures and modalities.
-
Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
A theoretical and empirical comparison claiming diffusion bridges have lower stochastic-optimal-control cost and greater robustness than flow matching when training data are scarce.
-
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
Three architecture-level interventions, overlapping convolutional tokenization plus sequence pooling, variadic attention receptive fields, and flow-aware distillation, each improve efficiency or quality in vision and ...
-
Deep Neural Networks Inspired by Differential Equations
A review of differential-equation-inspired neural networks that compiles known results into a taxonomy, with no new experiments or theory.
Discussion (0). Continue with ORCID to comment.