Pith. sign in

REVIEW 5 cited by

Inst-Inpaint: Instructing to Remove Objects with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03246 v2 pith:H2DQN3OO submitted 2023-04-06 cs.CV

classification cs.CV
keywords imageinpaintingobjectsremoveimagesinst-inpaintmasksmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image inpainting task refers to erasing unwanted pixels from images and filling them in a semantically consistent and realistic way. Traditionally, the pixels that are wished to be erased are defined with binary masks. From the application point of view, a user needs to generate the masks for the objects they would like to remove which can be time-consuming and prone to errors. In this work, we are interested in an image inpainting algorithm that estimates which object to be removed based on natural language input and removes it, simultaneously. For this purpose, first, we construct a dataset named GQA-Inpaint for this task. Second, we present a novel inpainting framework, Inst-Inpaint, that can remove objects from images based on the instructions given as text prompts. We set various GAN and diffusion-based baselines and run experiments on synthetic and real image datasets. We compare methods with different evaluation metrics that measure the quality and accuracy of the models and show significant quantitative and qualitative improvements.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  2. OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A single VLM-based reward model, conditionable on task and evaluation dimension, drives multi-task RL that improves mask-guided image editing across four tasks without task-specific SFT.

  3. Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional Inpainting

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A contact-aware, multi-regional inpainting method using a pretrained diffusion model completes occluded objects in human-object interactions without additional training.

  4. AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AttentionDrag is a one-step, training-free drag-editing method that uses diffusion self-attention to move regions, generate masks, and fill gaps.

  5. Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Subtracting the text embedding of an unwanted concept from a target word's embedding, with cross-attention keys and values steered in opposite directions, suppresses strongly entangled content in Stable Diffusion and ...

Pith tools