REVIEW 5 cited by
FateZero: Fusing Attentions for Zero-shot Text-based Video Editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The diffusion-based generative models have achieved remarkable success in text-based image generation. However, since it contains enormous randomness in generation progress, it is still challenging to apply such models for real-world visual content editing, especially in videos. In this paper, we propose FateZero, a zero-shot text-based editing method on real-world videos without per-prompt training or use-specific mask. To edit videos consistently, we propose several techniques based on the pre-trained models. Firstly, in contrast to the straightforward DDIM inversion technique, our approach captures intermediate attention maps during inversion, which effectively retain both structural and motion information. These maps are directly fused in the editing process rather than generated during denoising. To further minimize semantic leakage of the source video, we then fuse self-attentions with a blending mask obtained by cross-attention features from the source prompt. Furthermore, we have implemented a reform of the self-attention mechanism in denoising UNet by introducing spatial-temporal attention to ensure frame consistency. Yet succinct, our method is the first one to show the ability of zero-shot text-driven video style and local attribute editing from the trained text-to-image model. We also have a better zero-shot shape-aware editing ability based on the text-to-video model. Extensive experiments demonstrate our superior temporal consistency and editing capability than previous works.
Forward citations
Cited by 5 Pith papers
-
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking
TRACE anchors a video-diffusion editor to 3D meshes to perform consistent part-level edits on 3D Gaussian scenes in about 10 minutes per edit.
-
Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
A 2M-pair instruction-based video editing dataset built from real videos and specialist models, demonstrated to train editors that beat prior methods.
-
Low-Cost Test-Time Adaptation for Robust Video Editing
Vid-TTA proposes to adapt video editing UNets per test video via motion-aware masked autoencoding and prompt perturbation, with claimed but unquantified improvements.
-
Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models
ProLoRA transfers pre-trained LoRA, DoRA, and FouRA adapters between diffusion models in a single closed-form projection step, without retraining on the target model.
-
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
A visual heatmap tool lets people rate vision-language model reliability in video by inspecting patterns of green and red cells, with user ratings tracking objective F1 scores when those exist.
Discussion (0). Continue with ORCID to comment.