Pith. sign in

REVIEW 5 cited by

TurboEdit: Text-Based Image Editing Using Few-Step Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00735 v1 pith:S64XC46T submitted 2024-08-01 cs.CV cs.GR

classification cs.CVcs.GR
keywords editingtext-baseddiffusionartifactsimagenoiseapproachframeworks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models have opened the path to a wide range of text-based image editing frameworks. However, these typically build on the multi-step nature of the diffusion backwards process, and adapting them to distilled, fast-sampling methods has proven surprisingly challenging. Here, we focus on a popular line of text-based editing frameworks - the ``edit-friendly'' DDPM-noise inversion approach. We analyze its application to fast sampling methods and categorize its failures into two classes: the appearance of visual artifacts, and insufficient editing strength. We trace the artifacts to mismatched noise statistics between inverted noises and the expected noise schedule, and suggest a shifted noise schedule which corrects for this offset. To increase editing strength, we propose a pseudo-guidance approach that efficiently increases the magnitude of edits without introducing new artifacts. All in all, our method enables text-based image editing with as few as three diffusion steps, while providing novel insights into the mechanisms behind popular text-based editing approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instruction-based Image Manipulation by Watching How Things Move

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion editing model, InstructMove, is trained on video frame pairs annotated by MLLMs using spatial conditioning, enabling non-rigid edits and viewpoint changes.

  2. PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    PartEdit trains part-specific text tokens to produce spatial masks in a frozen diffusion model, enabling localized text-based part edits.

  3. ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A fine-tuned vision-language model converts images or text into JSON sewing-pattern descriptions, enabling estimation, generation, and editing of 3D garments.

  4. Stable Flow: Vital Layers for Training-Free Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.

  5. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

Pith tools