Pith. sign in

REVIEW 2 cited by

Diffusion Model-Based Video Editing: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07111 v1 pith:BUP476JJ submitted 2024-06-26 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords editingvideoapplicationsdiffusioncomprehensiveimageincludingmodel-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development of diffusion models (DMs) has significantly advanced image and video applications, making "what you want is what you see" a reality. Among these, video editing has gained substantial attention and seen a swift rise in research activity, necessitating a comprehensive and systematic review of the existing literature. This paper reviews diffusion model-based video editing techniques, including theoretical foundations and practical applications. We begin by overviewing the mathematical formulation and image domain's key methods. Subsequently, we categorize video editing approaches by the inherent connections of their core technologies, depicting evolutionary trajectory. This paper also dives into novel applications, including point-based editing and pose-guided human video editing. Additionally, we present a comprehensive comparison using our newly introduced V2VBench. Building on the progress achieved to date, the paper concludes with ongoing challenges and potential directions for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Integrating Large Language Models into Text Animation: An Intelligent Editing System with Inline and Chat Interaction

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A text-animation editor with inline and chat LLM agents was rated usable (SUS 75) by 11 non-professional testers.

  2. AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AdaFlow demonstrates a training-free method to edit more than 1,000 video frames in one inference on a single A800 GPU via adaptive attention token slimming and content-aware keyframe selection.

Pith tools