Pith. sign in

REVIEW 2 cited by

FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06269 v2 pith:RCMACQVH submitted 2024-03-10 cs.CV

classification cs.CV
keywords editingvideomodelsconsistencyefficientfastvideoeditgenerationspeed
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling in diffusion models poses challenges for efficient video editing. Existing approaches relying on image generation models for video editing suffer from time-consuming one-shot fine-tuning, additional condition extraction, or DDIM inversion, making real-time applications impractical. In this work, we propose FastVideoEdit, an efficient zero-shot video editing approach inspired by Consistency Models (CMs). By leveraging the self-consistency property of CMs, we eliminate the need for time-consuming inversion or additional condition extraction, reducing editing time. Our method enables direct mapping from source video to target video with strong preservation ability utilizing a special variance schedule. This results in improved speed advantages, as fewer sampling steps can be used while maintaining comparable generation quality. Experimental results validate the state-of-the-art performance and speed advantages of FastVideoEdit across evaluation metrics encompassing editing speed, temporal consistency, and text-video alignment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Visual Prompting for One-shot Controllable Video Editing without Inversion

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A one-shot video editing method that uses a 2x2 visual prompt grid, modified consistency sampling, and Stein Variational Gradient Descent to propagate first-frame edits without DDIM inversion.

  2. MoViE: Mobile Diffusion for Video Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MoViE distills a diffusion-based video editor into a single-step mobile model, achieving 12 fps on a Snapdragon 8 Gen 3 phone with modest quality loss.

Pith tools