Pith. sign in

REVIEW 10 cited by

Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.17599 v3 pith:ENZRLAK4 submitted 2023-03-30 cs.CV

classification cs.CV
keywords editingvideodiffusionimagemodelsmoduletrainingvid2vid-zero
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale text-to-image diffusion models achieve unprecedented success in image generation and editing. However, how to extend such success to video editing is unclear. Recent initial attempts at video editing require significant text-to-video data and computation resources for training, which is often not accessible. In this work, we propose vid2vid-zero, a simple yet effective method for zero-shot video editing. Our vid2vid-zero leverages off-the-shelf image diffusion models, and doesn't require training on any video. At the core of our method is a null-text inversion module for text-to-video alignment, a cross-frame modeling module for temporal consistency, and a spatial regularization module for fidelity to the original video. Without any training, we leverage the dynamic nature of the attention mechanism to enable bi-directional temporal modeling at test time. Experiments and analyses show promising results in editing attributes, subjects, places, etc., in real-world videos. Code is made available at \url{https://github.com/baaivision/vid2vid-zero}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free method preserves proxy-video dynamics via region-wise latent noising and Stochastic Flow Relaxation, outperforming editing and motion-transfer baselines on a custom 76-video set.

  2. AnchorSync: Global Consistency Optimization for Long Video Editing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    By jointly editing sparse anchor frames and interpolating with flow and edge guidance, AnchorSync produces temporally consistent edits on videos longer than previous diffusion methods could handle.

  3. Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An audio-conditioned video animation model is pretrained on noisy auto-curated videos and fine-tuned on a few clean examples, achieving top synchronization scores on a new 48-class benchmark with only 1.9% additional ...

  4. OutDreamer: Video Outpainting with a Diffusion Transformer

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OutDreamer couples a diffusion transformer with mask-driven self-attention and a latent alignment loss to outpaint videos in a zero-shot manner, exceeding prior zero-shot baselines on standard benchmarks.

  5. FADE: Frequency-Aware Diffusion Model Factorization for Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FADE edits real videos by guiding sampling with low-frequency attention-output differences from the first blocks of a text-to-video diffusion model, enabling training-free appearance and motion edits.

  6. AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AdaFlow demonstrates a training-free method to edit more than 1,000 video frames in one inference on a single A800 GPU via adaptive attention token slimming and content-aware keyframe selection.

  7. Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    EquiEdit balances temporal consistency and editability in diffusion-based text-guided video editing via a temporal Mamba module and spectral noise injection on initial latents.

  8. ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ObjFiller3D jointly optimizes a dense 360-degree ring of views to inpaint 3D objects with cross-view-consistent textures, reporting higher PSNR and LPIPS than per-view baselines at much lower runtime.

  9. Low-Cost Test-Time Adaptation for Robust Video Editing

    cs.CV 2025-07 reject novelty 5.0 of 10

    Vid-TTA proposes to adapt video editing UNets per test video via motion-aware masked autoencoding and prompt perturbation, with claimed but unquantified improvements.

  10. Light-A-Video: Training-free Video Relighting via Progressive Light Fusion

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Light-A-Video relights videos without training by injecting image relight results into a video diffusion model's denoising loop with cross-frame attention and progressive blending.

Pith tools