Pith. sign in

REVIEW 7 cited by

Dreamix: Video Diffusion Models are General Video Editors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.01329 v1 pith:F4OJUXDI submitted 2023-02-02 cs.CV

classification cs.CV
keywords videoimagediffusioneditinggeneralinformationmethodmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the first diffusion-based method that is able to perform text-based motion and appearance editing of general videos. Our approach uses a video diffusion model to combine, at inference time, the low-resolution spatio-temporal information from the original video with new, high resolution information that it synthesized to align with the guiding text prompt. As obtaining high-fidelity to the original video requires retaining some of its high-resolution information, we add a preliminary stage of finetuning the model on the original video, significantly boosting fidelity. We propose to improve motion editability by a new, mixed objective that jointly finetunes with full temporal attention and with temporal attention masking. We further introduce a new framework for image animation. We first transform the image into a coarse video by simple image processing operations such as replication and perspective geometric projections, and then use our general video editor to animate it. As a further application, we can use our method for subject-driven video generation. Extensive qualitative and numerical experiments showcase the remarkable editing ability of our method and establish its superior performance compared to baseline methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 43 citations worldwide. Full citation record

  1. ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Writing editing instructions that explicitly bind each attribute to a reference image via `<Image_N>` tokens substantially improves multi-reference video editing, and a specialized MLLM trained with GRPO generates the...

  2. ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free method preserves proxy-video dynamics via region-wise latent noising and Stochastic Flow Relaxation, outperforming editing and motion-transfer baselines on a custom 76-video set.

  3. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified model jointly segments, tracks, and captions all objects in videos, trained with VLM-generated synthetic captions, achieving SOTA on VidSTG, VLN, and BenSMOT.

  4. FADE: Frequency-Aware Diffusion Model Factorization for Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FADE edits real videos by guiding sampling with low-frequency attention-output differences from the first blocks of a text-to-video diffusion model, enabling training-free appearance and motion edits.

  5. MOVi: Training-free Text-conditioned Multi-Object Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MOVi improves multi-object video generation without retraining by using LLM-planned trajectories to reinitialize the diffusion noise and by reweighting attention to stop objects from mixing together.

  6. Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.

  7. TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    TV-LiVE injects key/value features from a source video into selected layers of a CogVideoX generator to perform object addition and non-rigid video editing without training.

Pith tools