Pith. sign in

REVIEW 1 cited by

RealCraft: Attention Control as A Tool for Zero-Shot Consistent Video Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12635 v4 pith:LOR32IWG submitted 2023-12-19 cs.CV

classification cs.CV
keywords editingvideovideoszero-shotacrossadditionalattentionattention-control-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Even though large-scale text-to-image generative models show promising performance in synthesizing high-quality images, applying these models directly to image editing remains a significant challenge. This challenge is further amplified in video editing due to the additional dimension of time. This is especially the case for editing real-world videos as it necessitates maintaining a stable structural layout across frames while executing localized edits without disrupting the existing content. In this paper, we propose RealCraft, an attention-control-based method for zero-shot real-world video editing. By swapping cross-attention for new feature injection and relaxing spatial-temporal attention of the editing object, we achieve localized shape-wise edit along with enhanced temporal consistency. Our model directly uses Stable Diffusion and operates without the need for additional information. We showcase the proposed zero-shot attention-control-based method across a range of videos, demonstrating shape-wise, time-consistent and parameter-free editing in videos of up to 64 frames.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DriveEditor uses depth-aware 3D bounding box projection and single-reference appearance cues to reposition, replace, remove, and insert objects in driving videos with a single diffusion framework.

Pith tools