Pith. sign in

REVIEW 4 cited by

BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06976 v1 pith:7GIT4I6P submitted 2024-03-11 cs.CV

classification cs.CV
keywords imageinpaintingbrushnetmaskedmodeladvancementsdiffusiondivision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image inpainting, the process of restoring corrupted images, has seen significant advancements with the advent of diffusion models (DMs). Despite these advancements, current DM adaptations for inpainting, which involve modifications to the sampling strategy or the development of inpainting-specific DMs, frequently suffer from semantic inconsistencies and reduced image quality. Addressing these challenges, our work introduces a novel paradigm: the division of masked image features and noisy latent into separate branches. This division dramatically diminishes the model's learning load, facilitating a nuanced incorporation of essential masked image information in a hierarchical fashion. Herein, we present BrushNet, a novel plug-and-play dual-branch model engineered to embed pixel-level masked image features into any pre-trained DM, guaranteeing coherent and enhanced image inpainting outcomes. Additionally, we introduce BrushData and BrushBench to facilitate segmentation-based inpainting training and performance assessment. Our extensive experimental analysis demonstrates BrushNet's superior performance over existing models across seven key metrics, including image quality, mask region preservation, and textual coherence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.

  2. DreamLight: Towards Harmonious and Consistent Image Relighting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A unified image- and text-based relighting model with direction-biased attention and a wavelet foreground fixer outperforms existing methods on a synthetic relighting benchmark.

  3. Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-step training loss plus a fine-tuned VAE reduces color and structure discontinuities at mask boundaries in diffusion image inpainting and outpainting.

  4. MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

    cs.CV 2025-06 conditional novelty 4.0 of 10

    MTADiffusion improves text-guided object inpainting by training on a new 5M-image mask-text dataset with edge prediction and style-consistency losses.

Pith tools