Pith. sign in

REVIEW 2 cited by

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.15260 v2 pith:SDGRJCKK submitted 2024-11-22 cs.CV cs.AI

classification cs.CVcs.AI
keywords editingvideovivid-10mdatasetlocalbaselinedatainteractive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle is the absence of open-source, large-scale video editing datasets based on real-world data, as constructing such datasets is both time-consuming and costly. Moreover, video data requires a significantly larger number of tokens for representation, which substantially increases the training costs for video editing models. Lastly, current video editing models offer limited interactivity, often making it difficult for users to express their editing requirements effectively in a single attempt. To address these challenges, this paper introduces a dataset VIVID-10M and a baseline model VIVID. VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, which comprises 9.7M samples that encompass a wide range of video editing tasks. VIVID is a Versatile and Interactive VIdeo local eDiting model trained on VIVID-10M, which supports entity addition, modification, and deletion. At its core, a keyframe-guided interactive video editing mechanism is proposed, enabling users to iteratively edit keyframes and propagate it to other frames, thereby reducing latency in achieving desired outcomes. Extensive experimental evaluations show that our approach achieves state-of-the-art performance in video local editing, surpassing baseline methods in both automated metrics and user studies. The VIVID-10M dataset are open-sourced at https://kwaivgi.github.io/VIVID/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage video object remover that removes text conditioning and uses minimax adversarial noise to achieve high-quality removal in 6 sampling steps without classifier-free guidance.

  2. Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A 2M-pair instruction-based video editing dataset built from real videos and specialist models, demonstrated to train editors that beat prior methods.

Pith tools