Pith. sign in

REVIEW 7 cited by

A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14555 v1 pith:FEVL7PJX submitted 2024-06-20 cs.CV

classification cs.CV
keywords editingimagemodelsdiffusionfieldframeworkscenarioscontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image editing aims to edit the given synthetic or real image to meet the specific requirements from users. It is widely studied in recent years as a promising and challenging field of Artificial Intelligence Generative Content (AIGC). Recent significant advancement in this field is based on the development of text-to-image (T2I) diffusion models, which generate images according to text prompts. These models demonstrate remarkable generative capabilities and have become widely used tools for image editing. T2I-based image editing methods significantly enhance editing performance and offer a user-friendly interface for modifying content guided by multimodal inputs. In this survey, we provide a comprehensive review of multimodal-guided image editing techniques that leverage T2I diffusion models. First, we define the scope of image editing from a holistic perspective and detail various control signals and editing scenarios. We then propose a unified framework to formalize the editing process, categorizing it into two primary algorithm families. This framework offers a design space for users to achieve specific goals. Subsequently, we present an in-depth analysis of each component within this framework, examining the characteristics and applicable scenarios of different combinations. Given that training-based methods learn to directly map the source image to target one under user guidance, we discuss them separately, and introduce injection schemes of source image in different scenarios. Additionally, we review the application of 2D techniques to video editing, highlighting solutions for inter-frame inconsistency. Finally, we discuss open challenges in the field and suggest potential future research directions. We keep tracing related works at https://github.com/xinchengshuai/Awesome-Image-Editing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Appearance pointers are compact tokens that let a diffusion transformer apply text, image, or combined prompts to specific image regions in a single pass.

  2. CharaConsist: Fine-Grained Consistent Character Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A training-free consistency method for text-to-image DiT models that keeps characters and backgrounds stable across shots via point-tracking attention, adaptive token merge, and foreground/background masking.

  3. AnyI2V: Animating Any Conditional Image with Motion Control

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.

  4. MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    MMAFFBen is an open-source multilingual and multimodal benchmark for evaluating sentiment and emotion understanding of LLMs and VLMs.

  5. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.

  6. Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting editing pipeline that combines 2D diffusion localization, depth-based point seeding, and sequential view refinement to achieve up to 4x faster local edits.

  7. A Survey on Pre-Trained Diffusion Model Distillations

    cs.LG 2025-02 unverdicted novelty 2.0 of 10

    A taxonomy of pre-trained diffusion model distillation methods grouped into fidelity, trajectory, and adversarial losses.

Pith tools