Pith. sign in

REVIEW 5 cited by

Magic Insert: Style-Aware Drag-and-Drop

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02489 v1 pith:G7IW4YIQ submitted 2024-07-02 cs.CV cs.AIcs.GRcs.HCcs.LG

classification cs.CVcs.AIcs.GRcs.HCcs.LG
keywords imagemethodstyle-awareinsertionobjectstyletargetdomain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the problem of style-aware drag-and-drop and presents a method for tackling it by addressing two sub-problems: style-aware personalization and realistic object insertion in stylized images. For style-aware personalization, our method first fine-tunes a pretrained text-to-image diffusion model using LoRA and learned text tokens on the subject image, and then infuses it with a CLIP representation of the target style. For object insertion, we use Bootstrapped Domain Adaption to adapt a domain-specific photorealistic object insertion model to the domain of diverse artistic styles. Overall, the method significantly outperforms traditional approaches such as inpainting. Finally, we present a dataset, SubjectPlop, to facilitate evaluation and future progress in this area. Project page: https://magicinsert.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Raw CSD cosine similarity produces negative discrimination gaps for many artists and does not support absolute style-fidelity interpretation, but CSLS readout on frozen backbones reduces failures and improves AUC.

  2. Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single reference image guides a diffusion model to insert coherent objects into camera-plus-lidar driving scenes and to insert mammographic anomalies into new scans.

  3. AIComposer: Any Style and Content Image Composition via Feature Integration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A nearly training-free SDXL pipeline composes foreground content with background style using a small MLP that merges CLIP image features, removing the need for text prompts.

  4. HOComp: Interaction-Aware Human-Object Composition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-transformer method that composes a foreground object into a human image with MLLM-chosen interaction regions, pose keypoint supervision, and appearance/background consistency losses, plus a new paired dataset.

  5. Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting editing pipeline that combines 2D diffusion localization, depth-based point seeding, and sequential view refinement to achieve up to 4x faster local edits.

Pith tools