Pith. sign in

REVIEW 2 cited by

FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18071 v1 pith:I6GNSNXS submitted 2024-09-26 cs.CV cs.AI

classification cs.CVcs.AI
keywords editingimagefreeeditreferenceinstructionsfreebenchlanguagereference-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based image editing, which can accurately reproduce the visual concept from the reference image based on user-friendly language instructions. Our approach leverages the multi-modal instruction encoder to encode language instructions to guide the editing process. This implicit way of locating the editing area eliminates the need for manual editing masks. To enhance the reconstruction of reference details, we introduce the Decoupled Residual ReferAttention (DRRA) module. This module is designed to integrate fine-grained reference features extracted by a detail extractor into the image editing process in a residual way without interfering with the original self-attention. Given that existing datasets are unsuitable for reference-based image editing tasks, particularly due to the difficulty in constructing image triplets that include a reference image, we curate a high-quality dataset, FreeBench, using a newly developed twice-repainting scheme. FreeBench comprises the images before and after editing, detailed editing instructions, as well as a reference image that maintains the identity of the edited object, encompassing tasks such as object addition, replacement, and deletion. By conducting phased training on FreeBench followed by quality tuning, FreeEdit achieves high-quality zero-shot editing through convenient language instructions. We conduct extensive experiments to evaluate the effectiveness of FreeEdit across multiple task types, demonstrating its superiority over existing methods. The code will be available at: https://freeedit.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AvatarMakeup: Realistic Makeup Transfer for 3D Animatable Head Avatars

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A coarse-to-fine pipeline transfers makeup from one reference image to an animatable 3D Gaussian avatar, using UV-map averaging for cross-view consistency and diffusion refinement for detail.

  2. Borrowing from anything: A generalizable framework for reference-guided instance editing

    cs.CV 2025-12 conditional novelty 4.0 of 10

    GENIE uses spatial alignment, residual feature scaling, and progressive attention fusion to transfer a reference's appearance onto a target, achieving state-of-the-art scores on AnyInsertion.

Pith tools