REVIEW 14 cited by
FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Editing real images using a pre-trained text-to-image (T2I) diffusion/flow model often involves inverting the image into its corresponding noise map. However, inversion by itself is typically insufficient for obtaining satisfactory results, and therefore many methods additionally intervene in the sampling process. Such methods achieve improved results but are not seamlessly transferable between model architectures. Here, we introduce FlowEdit, a text-based editing method for pre-trained T2I flow models, which is inversion-free, optimization-free and model agnostic. Our method constructs an ODE that directly maps between the source and target distributions (corresponding to the source and target text prompts) and achieves a lower transport cost than the inversion approach. This leads to state-of-the-art results, as we illustrate with Stable Diffusion 3 and FLUX. Code and examples are available on the project's webpage.
Forward citations
Cited by 14 Pith papers
-
Feedforward 3D Editing Learns from Semantic-Part Transformation
Semantic-part-grounded paired data (Pxform) plus a source-controlled feedforward editor (PartFlow) yields state-of-the-art geometry and appearance 3D editing without inference-time 3D masks.
-
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
A large human-annotated benchmark of AI-edited images (EBench-18K) plus a fine-tuned LMM metric (LMM4Edit) that predicts human preference scores across three dimensions and answers editing-specific questions.
-
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
ReFlex edits real images with FLUX by extracting attention and residual features from a mid-step latent and adapting them during generation, improving text alignment while preserving structure.
-
DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
Direct Noise Alignment iteratively moves a random Gaussian noise until the model's predicted velocity matches the straight-line velocity to the image, reducing inversion drift and giving the best reported fidelity-edi...
-
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.
-
SemanticAudio: Audio Generation and Editing in Semantic Space
SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...
-
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
A grid-based LoRA training scheme lets a text-to-video model personalize previously unseen dynamic concepts, subject appearance plus motion, in one feedforward pass with no per-video fine-tuning.
-
FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
FreeCus is a training-free method that combines pivotal attention sharing, reversed noise shifting, and MLLM captions to personalize Flux.1 text-to-image generation from a single reference image.
-
Minimalist Concept Erasure in Generative Models
A final-output-only loss with learned neuron masks erases concepts from flow-based image generators more robustly than per-step fine-tuning methods.
-
EditP23: 3D Editing via Propagation of Image Prompts to Multi-View
A training-free, mask-free 3D editing method that propagates a single user-edited 2D view across all views of a frozen multi-view diffusion model.
-
Translationese as a Rational Response to Translation Task Difficulty
Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.
-
FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
FlowAlign adds a terminal-point source-similarity regularization to inversion-free flow-based editing, improving structural consistency while maintaining semantic alignment.
-
ACE-Step: A Step Towards Music Generation Foundation Model
ACE-Step is a fast, controllable open-source music generation model built from a mel-spectrogram DCAE, a linear DiT, and REPA-style semantic alignment.
-
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.
Discussion (0). Continue with ORCID to comment.