Pith. sign in

REVIEW 15 cited by

AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14468 v4 pith:M55BPJF3 submitted 2024-03-21 cs.CV cs.AIcs.MM

classification cs.CVcs.AIcs.MM
keywords editingvideoanyv2veditsmethodsmodelstasksexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the dynamic field of digital content creation using generative models, state-of-the-art video editing models still do not offer the level of quality and control that users desire. Previous works on video editing either extended from image-based generative models in a zero-shot manner or necessitated extensive fine-tuning, which can hinder the production of fluid video edits. Furthermore, these methods frequently rely on textual input as the editing guidance, leading to ambiguities and limiting the types of edits they can perform. Recognizing these challenges, we introduce AnyV2V, a novel tuning-free paradigm designed to simplify video editing into two primary steps: (1) employing an off-the-shelf image editing model to modify the first frame, (2) utilizing an existing image-to-video generation model to generate the edited video through temporal feature injection. AnyV2V can leverage any existing image editing tools to support an extensive array of video editing tasks, including prompt-based editing, reference-based style transfer, subject-driven editing, and identity manipulation, which were unattainable by previous methods. AnyV2V can also support any video length. Our evaluation shows that AnyV2V achieved CLIP-scores comparable to other baseline methods. Furthermore, AnyV2V significantly outperformed these baselines in human evaluations, demonstrating notable improvements in visual consistency with the source video while producing high-quality edits across all editing tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single video-diffusion framework composites both static images and dynamic footage along user-defined trajectories by transporting canonical foreground latents directly into the background latent sequence.

  2. GraphVid: Interactive Graph-Controllable Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GraphVid controls video generation with user-editable interaction scene graphs, reporting FID/FVD improvements over trajectory- and text-physics baselines using 0.6B trainable parameters.

  3. TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

    cs.CV 2026-04 conditional novelty 6.0 of 10

    TRACE anchors a video-diffusion editor to 3D meshes to perform consistent part-level edits on 3D Gaussian scenes in about 10 minutes per edit.

  4. InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

    cs.CV 2025-12 conditional novelty 6.0 of 10

    InsertAnywhere inserts a reference object into arbitrary videos by reconstructing 4D geometry to propagate a user-given placement across frames and fine-tuning video diffusion on ROSE++, a removal-to-insertion dataset...

  5. UniVideo: Unified Understanding, Generation, and Editing for Videos

    cs.CV 2025-10 conditional novelty 6.0 of 10

    UniVideo combines a frozen MLLM and a video DiT to unify video understanding, generation, in-context editing, visual prompting, and zero-shot free-form video edits under one instruction interface.

  6. AnchorSync: Global Consistency Optimization for Long Video Editing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    By jointly editing sparse anchor frames and interpolating with flow and edge guidance, AnchorSync produces temporally consistent edits on videos longer than previous diffusion methods could handle.

  7. Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EVS combines a text-to-image and a text-to-video diffusion model in a single denoising pass, improving frame quality and temporal consistency without retraining.

  8. Populate-A-Scene: Affordance-Aware Human Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A fine-tuned text-to-video model inserts a person into a scene and generates an interaction video without bounding boxes or pose input, and its attention maps reveal a latent sense of affordance.

  9. UNIC: Unified In-Context Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.

  10. IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

    cs.CV 2025-06 reject novelty 6.0 of 10

    A diffusion video model that jointly uses HDR lighting, relit frames, and 3D point tracks to relight videos from text prompts.

  11. MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage video object remover that removes text conditioning and uses minimax adversarial noise to achieve high-quality removal in 6 sampling steps without classifier-free guidance.

  12. Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A reference-based video editing pipeline that guides cross-image attention with diffusion correspondence, then trains a per-video restoration model to clean up the zero-shot output.

  13. On the Astrophysical Origin of Binary Black Hole Subpopulations: A Tale of Three Channels?

    astro-ph.HE 2026-03 unverdicted novelty 5.0 of 10

    Parametrized mixture models of LIGO-Virgo-KAGRA BBHs favor three channels—isolated binaries (~79%), globular-cluster dynamics (~14.5%), and higher-generation mergers (~2.5%)—with fractions evolving in redshift.

  14. WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A bidirectional egocentric-to-exocentric video translation framework trained with in-context attention on a new synthetic+real dataset, with evaluation flaws around reference leakage and missing direct baselines.

  15. DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.

Pith tools