REVIEW 14 cited by
Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-guided diffusion models have revolutionized image generation and editing, offering exceptional realism and diversity. Specifically, in the context of diffusion-based editing, where a source image is edited according to a target prompt, the process commences by acquiring a noisy latent vector corresponding to the source image via the diffusion model. This vector is subsequently fed into separate source and target diffusion branches for editing. The accuracy of this inversion process significantly impacts the final editing outcome, influencing both essential content preservation of the source image and edit fidelity according to the target prompt. Prior inversion techniques aimed at finding a unified solution in both the source and target diffusion branches. However, our theoretical and empirical analyses reveal that disentangling these branches leads to a distinct separation of responsibilities for preserving essential content and ensuring edit fidelity. Building on this insight, we introduce "Direct Inversion," a novel technique achieving optimal performance of both branches with just three lines of code. To assess image editing performance, we present PIE-Bench, an editing benchmark with 700 images showcasing diverse scenes and editing types, accompanied by versatile annotations and comprehensive evaluation metrics. Compared to state-of-the-art optimization-based inversion techniques, our solution not only yields superior performance across 8 editing methods but also achieves nearly an order of speed-up.
Forward citations
Cited by 14 Pith papers
-
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
ReFlex edits real images with FLUX by extracting attention and residual features from a mid-step latent and adapting them during generation, improving text alignment while preserving structure.
-
DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
Direct Noise Alignment iteratively moves a random Gaussian noise until the model's predicted velocity matches the straight-line velocity to the image, reducing inversion drift and giving the best reported fidelity-edi...
-
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.
-
FARI: Robust One-Step Inversion for Watermarking in Diffusion Models
One-step adversarially LoRA-tuned inversion exploits low-curvature reverse trajectories to beat 50-step DDIM on watermark robustness after ~20 minutes of fine-tuning.
-
Making Implicit Preservation Intent Explicit in Conversational Image Editing
Conversational image editors fail to restore temporarily occluded content; ReSpec fixes this by explicitly selecting historical visual references and rewriting instructions to guide restoration.
-
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
VARIN uses a Location-aware Argmax Inversion pseudo-inverse of Gumbel-max sampling to extract editable discrete noises, enabling training-free prompt-guided editing for visual autoregressive models.
-
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...
-
Trade-offs in Image Generation: How Do Different Dimensions Interact?
A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.
-
AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
AttentionDrag is a one-step, training-free drag-editing method that uses diffusion self-attention to move regions, generate masks, and fill gaps.
-
Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing
A region and semantic aware metric for instruction-based image editing, built from LLM parsing plus detection, segmentation, and CLIP directional similarity, reports the highest human alignment among compared metrics.
-
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
RefEdit-Bench measures referring-expression image editing; the RefEdit model, trained on 20K synthetic triplets, reports state-of-the-art results over million-scale baselines.
-
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.
-
FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
FlowAlign adds a terminal-point source-similarity regularization to inversion-free flow-based editing, improving structural consistency while maintaining semantic alignment.
-
DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing
DCI combines reference-guided noise correction with fixed-point latent refinement and reports state-of-the-art reconstruction and editing metrics on PIE-Bench.
Discussion (0). Sign in to comment.