Pith. sign in

REVIEW 5 cited by

FateZero: Fusing Attentions for Zero-shot Text-based Video Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.09535 v3 pith:YN5PNLDE submitted 2023-03-16 cs.CV

classification cs.CV
keywords editingzero-shotmodelstext-basedvideovideosabilityattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The diffusion-based generative models have achieved remarkable success in text-based image generation. However, since it contains enormous randomness in generation progress, it is still challenging to apply such models for real-world visual content editing, especially in videos. In this paper, we propose FateZero, a zero-shot text-based editing method on real-world videos without per-prompt training or use-specific mask. To edit videos consistently, we propose several techniques based on the pre-trained models. Firstly, in contrast to the straightforward DDIM inversion technique, our approach captures intermediate attention maps during inversion, which effectively retain both structural and motion information. These maps are directly fused in the editing process rather than generated during denoising. To further minimize semantic leakage of the source video, we then fuse self-attentions with a blending mask obtained by cross-attention features from the source prompt. Furthermore, we have implemented a reform of the self-attention mechanism in denoising UNet by introducing spatial-temporal attention to ensure frame consistency. Yet succinct, our method is the first one to show the ability of zero-shot text-driven video style and local attribute editing from the trained text-to-image model. We also have a better zero-shot shape-aware editing ability based on the text-to-video model. Extensive experiments demonstrate our superior temporal consistency and editing capability than previous works.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

    cs.CV 2026-04 conditional novelty 6.0 of 10

    TRACE anchors a video-diffusion editor to 3D meshes to perform consistent part-level edits on 3D Gaussian scenes in about 10 minutes per edit.

  2. Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A 2M-pair instruction-based video editing dataset built from real videos and specialist models, demonstrated to train editors that beat prior methods.

  3. Low-Cost Test-Time Adaptation for Robust Video Editing

    cs.CV 2025-07 reject novelty 5.0 of 10

    Vid-TTA proposes to adapt video editing UNets per test video via motion-aware masked autoencoding and prompt perturbation, with claimed but unquantified improvements.

  4. Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models

    cs.AI 2025-05 conditional novelty 5.0 of 10

    ProLoRA transfers pre-trained LoRA, DoRA, and FouRA adapters between diffusion models in a single closed-form projection step, without retraining on the target model.

  5. IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A visual heatmap tool lets people rate vision-language model reliability in video by inspecting patterns of green and red cells, with user ratings tracking objective F1 scores when those exist.

Pith tools