Pith. sign in

REVIEW 4 cited by

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13185 v4 pith:RP2WMC4U submitted 2024-02-20 cs.CV

classification cs.CV
keywords editingvideomotionappearanceframeworktemporalunieditfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.g., from eating to waving), which distinguishes video editing from image editing, is underexplored. In this work, we present UniEdit, a tuning-free framework that supports both video motion and appearance editing by harnessing the power of a pre-trained text-to-video generator within an inversion-then-generation framework. To realize motion editing while preserving source video content, based on the insights that temporal and spatial self-attention layers encode inter-frame and intra-frame dependency respectively, we introduce auxiliary motion-reference and reconstruction branches to produce text-guided motion and source features respectively. The obtained features are then injected into the main editing path via temporal and spatial self-attention layers. Extensive experiments demonstrate that UniEdit covers video motion editing and various appearance editing scenarios, and surpasses the state-of-the-art methods. Our code will be publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LayerFlow: A Unified Model for Layer-aware Video Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    LayerFlow is a unified diffusion-transformer model that generates transparent foreground, background, and blended video layers from per-layer prompts, and supports decomposition and conditioned generation in one framework.

  2. HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.

  3. STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STR-Match edits videos without retraining by matching source and target 'spatiotemporal relevance scores' derived from attention maps during latent optimization.

  4. Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A review that organizes camera trajectory generation into representation levels, algorithm families, evaluation metrics, and datasets.

Pith tools