REVIEW 19 cited by
ControlVideo: Training-free Controllable Text-to-Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temporal modeling. Besides the training burden, the generated videos also suffer from appearance inconsistency and structural flickers, especially in long video synthesis. To address these challenges, we design a \emph{training-free} framework called \textbf{ControlVideo} to enable natural and efficient text-to-video generation. ControlVideo, adapted from ControlNet, leverages coarsely structural consistency from input motion sequences, and introduces three modules to improve video generation. Firstly, to ensure appearance coherence between frames, ControlVideo adds fully cross-frame interaction in self-attention modules. Secondly, to mitigate the flicker effect, it introduces an interleaved-frame smoother that employs frame interpolation on alternated frames. Finally, to produce long videos efficiently, it utilizes a hierarchical sampler that separately synthesizes each short clip with holistic coherency. Empowered with these modules, ControlVideo outperforms the state-of-the-arts on extensive motion-prompt pairs quantitatively and qualitatively. Notably, thanks to the efficient designs, it generates both short and long videos within several minutes using one NVIDIA 2080Ti. Code is available at https://github.com/YBYBZhang/ControlVideo.
Forward citations
Cited by 19 Pith papers
-
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
EmoWorld adds three training-free steering operators to a frozen video diffusion transformer that separately control atmosphere, affect-bearing cues, and temporal emotion transitions in generated videos.
-
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
SGC quantifies 3D geometric consistency of generated videos by measuring divergence among local camera poses estimated only on static background sub-regions.
-
ANYPORTAL: Zero-Shot Consistent Video Background Replacement
A training-free video background replacement pipeline that keeps the foreground pixel-consistent by projecting refined latents through a deterministic reparameterization.
-
DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.
-
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.
-
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
EVS combines a text-to-image and a text-to-video diffusion model in a single denoising pass, improving frame quality and temporal consistency without retraining.
-
HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
HairShifter transfers a reference hairstyle onto a person throughout a video by animating a high-quality anchor frame and using a gated decoder that preserves non-hair regions.
-
AnyI2V: Animating Any Conditional Image with Motion Control
AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.
-
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
MotionEcho adaptively re-injects teacher-model guidance into few-step distilled video generators so reference motion can be copied at test time without training.
-
Controllable Coupled Image Generation via Diffusion Models
A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and tex...
-
UNIC: Unified In-Context Video Editing
One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.
-
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
By training a semantic expert and a LoRA-based detail expert, DCM reaches nearly teacher-level VBench scores with 4-step video sampling on HunyuanVideo and CogVideoX.
-
Interactive Video Generation via Domain Adaptation
A training-free method combines mask normalization and temporal intrinsic denoising to improve trajectory control and perceptual quality in text-to-video diffusion.
-
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.
-
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.
-
Multi-View Face and Gesture Animation with Dynamic Gaussians
Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.
-
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
LongVie combines unified noise initialization, global control normalization, and multi-modal depth-plus-keypoint guidance to generate temporally consistent controllable videos of up to one minute.
-
EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control
A training-free Lagrangian motion magnification framework with periodic reference resetting and tissue-aware dual-mask control improves vascular pulsation visibility in endoscopic surgery videos.
-
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.
Discussion (0). Sign in to comment.