REVIEW 7 cited by
Dreamix: Video Diffusion Models are General Video Editors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the first diffusion-based method that is able to perform text-based motion and appearance editing of general videos. Our approach uses a video diffusion model to combine, at inference time, the low-resolution spatio-temporal information from the original video with new, high resolution information that it synthesized to align with the guiding text prompt. As obtaining high-fidelity to the original video requires retaining some of its high-resolution information, we add a preliminary stage of finetuning the model on the original video, significantly boosting fidelity. We propose to improve motion editability by a new, mixed objective that jointly finetunes with full temporal attention and with temporal attention masking. We further introduce a new framework for image animation. We first transform the image into a coarse video by simple image processing operations such as replication and perspective geometric projections, and then use our general video editor to animate it. As a further application, we can use our method for subject-driven video generation. Extensive qualitative and numerical experiments showcase the remarkable editing ability of our method and establish its superior performance compared to baseline methods.
Forward citations
Cited by 7 Pith papers
-
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
Writing editing instructions that explicitly bind each attribute to a reference image via `<Image_N>` tokens substantially improves multi-reference video editing, and a specialized MLLM trained with GRPO generates the...
-
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
A training-free method preserves proxy-video dynamics via region-wise latent noising and Stochastic Flow Relaxation, outperforming editing and motion-transfer baselines on a custom 76-video set.
-
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
A unified model jointly segments, tracks, and captions all objects in videos, trained with VLM-generated synthetic captions, achieving SOTA on VidSTG, VLN, and BenSMOT.
-
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
FADE edits real videos by guiding sampling with low-frequency attention-output differences from the first blocks of a text-to-video diffusion model, enabling training-free appearance and motion edits.
-
MOVi: Training-free Text-conditioned Multi-Object Video Generation
MOVi improves multi-object video generation without retraining by using LLM-planned trajectories to reinitialize the diffusion noise and by reweighting attention to stop objects from mixing together.
-
Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization
ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.
-
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
TV-LiVE injects key/value features from a source video into selected layers of a CogVideoX generator to perform object addition and non-rigid video editing without training.
Discussion (0). Continue with ORCID to comment.