Pith. sign in

REVIEW 14 cited by

FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.08629 v2 pith:IFGB4JQE submitted 2024-12-11 cs.CV cs.LG

classification cs.CVcs.LG
keywords editingflowmodelpre-trainedresultscorrespondingdiffusionflowedit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Editing real images using a pre-trained text-to-image (T2I) diffusion/flow model often involves inverting the image into its corresponding noise map. However, inversion by itself is typically insufficient for obtaining satisfactory results, and therefore many methods additionally intervene in the sampling process. Such methods achieve improved results but are not seamlessly transferable between model architectures. Here, we introduce FlowEdit, a text-based editing method for pre-trained T2I flow models, which is inversion-free, optimization-free and model agnostic. Our method constructs an ODE that directly maps between the source and target distributions (corresponding to the source and target text prompts) and achieves a lower transport cost than the inversion approach. This leads to state-of-the-art results, as we illustrate with Stable Diffusion 3 and FLUX. Code and examples are available on the project's webpage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feedforward 3D Editing Learns from Semantic-Part Transformation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Semantic-part-grounded paired data (Pxform) plus a source-controlled feedforward editor (PartFlow) yields state-of-the-art geometry and appearance 3D editing without inference-time 3D masks.

  2. LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A large human-annotated benchmark of AI-edited images (EBench-18K) plus a fine-tuned LMM metric (LMM4Edit) that predicts human preference scores across three dimensions and answers editing-specific questions.

  3. ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    ReFlex edits real images with FLUX by extracting attention and residual features from a mid-step latent and adapting them during generation, improving text alignment while preserving structure.

  4. DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Direct Noise Alignment iteratively moves a random Gaussian noise until the model's predicted velocity matches the straight-line velocity to the image, reducing inversion drift and giving the best reported fidelity-edi...

  5. EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.

  6. SemanticAudio: Audio Generation and Editing in Semantic Space

    eess.AS 2026-01 conditional novelty 6.0 of 10

    SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...

  7. Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A grid-based LoRA training scheme lets a text-to-video model personalize previously unseen dynamic concepts, subject appearance plus motion, in one feedforward pass with no per-video fine-tuning.

  8. FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FreeCus is a training-free method that combines pivotal attention sharing, reversed noise shifting, and MLLM captions to personalize Flux.1 text-to-image generation from a single reference image.

  9. Minimalist Concept Erasure in Generative Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A final-output-only loss with learned neuron masks erases concepts from flow-based image generators more robustly than per-step fine-tuning methods.

  10. EditP23: 3D Editing via Propagation of Image Prompts to Multi-View

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A training-free, mask-free 3D editing method that propagates a single user-edited 2D view across all views of a frozen multi-view diffusion model.

  11. Translationese as a Rational Response to Translation Task Difficulty

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.

  12. FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing

    cs.CV 2025-05 conditional novelty 5.0 of 10

    FlowAlign adds a terminal-point source-similarity regularization to inversion-free flow-based editing, improving structural consistency while maintaining semantic alignment.

  13. ACE-Step: A Step Towards Music Generation Foundation Model

    cs.SD 2025-05 conditional novelty 5.0 of 10

    ACE-Step is a fast, controllable open-source music generation model built from a mel-spectrogram DCAE, a linear DiT, and REPA-style semantic alignment.

  14. DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.

Pith tools