Pith. sign in

REVIEW 8 cited by

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08465 v1 pith:R2ZAVY6T submitted 2023-04-17 cs.CV

classification cs.CV
keywords editingimagegenerationconsistentmasactrlself-attentioncomplexexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex non-rigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Drag-based editing becomes pixel-space bidirectional warping plus inpainting, giving real-time previews and 0.3s final edits at 512x512.

  4. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  5. MARBLE: Material Recomposition and Blending in CLIP-Space

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.

  6. Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling

    cs.CV 2026-02 reject novelty 5.0 of 10

    A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.

  7. FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

    cs.GR 2025-07 conditional novelty 5.0 of 10

    FlowDrag combines 3D mesh deformation with diffusion-based drag editing, using the resulting 2D vector flow to steer the denoising process, and adds a ground-truth benchmark built from video frames.

  8. DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing

    cs.CV 2025-06 reject novelty 4.0 of 10

    DCI combines reference-guided noise correction with fixed-point latent refinement and reports state-of-the-art reconstruction and editing metrics on PIE-Bench.

Pith tools