REVIEW 5 major objections 6 minor 64 references
MotionPro: A Precise Motion Controller for Image-to-Video Generation
T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read MotionPro conditions image-to-video diffusion on trajectories sampled in local regions plus a motion mask, yielding precise control that separates object motion from camera motion.
desk verdict A plausible I2V control recipe with a useful new benchmark, but the headline WebVid FVD gap is not a clean measurement and the evaluation needs more rigor. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The control representation is the pair consisting of a region-wise trajectory and a motion mask. The region-wise trajectory is sampled from dense optical flow after masking: flow maps and visibility masks come from a tracking model, a global visibility mask is their frame-wise intersection, and the masked flow is split into $k\times k$ regions ($k=8$) from which trajectories are kept with a random mask ratio in $[0.95, 1.0]$. The motion mask is the per-pixel average flow magnitude across frames thresholded at 1, repeated to video length, marking the moving region so object motion is not confused with camera motion. These conditions feed a motion encoder whose multi-scale features predict scale and bias used to modulate group-normalized latent features in the diffusion UNet (adaptive feature modulation, with temporal convolutions zero-initialized), and low-rank adaptation layers fine-tune all attention modules to align generated motion with the input trajectory.
What would settle it
Train MotionPro twice on the same videos: once with the automatic tracker-extracted conditions, and once with conditions that simulate real user input (a single trajectory point padded to the local region and an annotator-drawn mask), then compare MD-Img and MD-Vid on a held-out set of user-annotated trajectories. If the second run does not preserve the reported advantage over the flow-densification baseline, the transfer assumption is false.
Extended reading notes
Core claim
The central discovery is that replacing Gaussian-filtered trajectory conditions with region-wise trajectories plus a motion mask makes image-to-video motion control both finer and more semantically correct. Training conditions are generated automatically: a dense optical tracker produces optical flow maps and visibility masks for each video; a global visibility mask is computed as their intersection, the masked flow maps are split into 8-by-8 local regions, and trajectories are kept in a randomly sampled subset (mask ratio between 0.95 and 1.0) to mimic sparse user input. The motion mask is built by thresholding the average flow magnitude at 1 and repeating it over frames. The two signals are concatenated, encoded into multi-scale features, and injected into a pretrained video diffusion UNet by predicting scale and bias that modulate group-normalized latent features, while low-rank adaptation layers fine-tune every attention module to improve trajectory alignment. The paper's quantitative anchor is a Frechet Video Distance (FVD) of 59.88 on the large benchmark versus 87.70 for the best prior approach, together with lower values on two mean-distance trajectory-alignment metrics (MD-Img and MD-Vid) on the new user-annotated benchmark for both fine-grained and object-level control.
Load-bearing premise
Training conditions are extracted automatically by a dense optical tracker from real videos, and the method assumes these flow-derived trajectories and thresholded masks are statistically representative of the sparse single-point trajectories and brushed masks users actually draw; if user inputs live outside the tracker's flow distribution, the measured gains would not transfer to interactive use.
Editorial extensions
If this is right
- The method reports FVD 59.88 on a large public video benchmark, outperforming the best prior approach by 27.82, with better FID and comparable frame consistency.
- On the new user-annotated benchmark, it achieves lower MD-Img and MD-Vid distances for both fine-grained local-part motion and object-level motion, indicating better alignment between drawn trajectories and generated video motion.
- Integrating the motion mask as a conditional input, rather than only as post-processing flow masking, reduces the misinterpretation of object motion as camera motion.
- The learned controller transfers to camera control without retraining by converting camera poses into sparse trajectories and using an all-ones motion mask.
- Ablations identify local region size 8 and minimal mask ratio 0.95 as the best trade-off; larger regions blur fine control, smaller regions weaken the conditioning signal.
Reading between the lines
- Because the motion mask is robust to rough shapes, semantic or click-based masks could replace the brushed mask at inference with minimal retraining; this is an extension the paper motivates but does not test.
- The random mask ratio in [0.95, 1.0] acts as a dropout on the training condition; ablating only the ratio distribution would isolate how much of the robustness comes from this augmentation.
- The same two-signal conditioning could be applied to trajectory-based video editing and long-video generation, where object/camera ambiguity is more severe and currently resolved by hand.
- Because the paper reports a higher average flow magnitude for its generated videos than the densification baseline, part of the FVD gain may reflect greater motion diversity rather than pure trajectory accuracy; separating the two would clarify the source of the improvement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MotionPro, an image-to-video (I2V) generation controller built on Stable Video Diffusion. Instead of conditioning on Gaussian-filtered point trajectories, MotionPro uses region-wise trajectories sampled in k×k local regions plus a motion mask derived from dense optical flow. These signals are injected into the frozen 3D-UNet via adaptive feature modulation, and all attention modules are fine-tuned with LoRA. The paper also introduces MC-Bench, a new benchmark of 1.1K user-annotated image-trajectory pairs, and reports improvements over DragNUWA, MOFA-Video, DragAnything, and DragDiffusion on WebVid-10M and MC-Bench in terms of FVD, FID, frame consistency, and trajectory-alignment metrics.
Significance. If the central claims hold, the paper offers a useful and simple conditioning representation: region-wise trajectories plus a motion mask, which avoids the coarse smoothing of large Gaussian kernels and gives a natural way to distinguish object motion from camera motion. The design is straightforward to implement on top of SVD, and the proposed MC-Bench could be valuable to the community if released. However, the reported evidence has important gaps. The main WebVid-10M FVD comparison appears to be confounded by the amount of motion information given to each method; hyperparameters are selected on and then evaluated on the same MC-Bench; and all quantitative results are single-run point estimates without uncertainty. These issues affect the strength of the empirical support for the representation claim, though they do not invalidate the core idea.
major comments (5)
- [§4.1 and §3.4, Table 1] The WebVid-10M comparison is confounded by condition density. Section 4.1 says trajectories are sampled at a ratio of 15%, while Section 3.4 says a user trajectory is padded to the surrounding k×k region (k=8). If the 15%-sampled points are also padded during evaluation, then for a 320×512 frame the 15% of pixels (~24.6K points) each contribute a 64-pixel block, giving roughly 1.57M pixel-coverage attempts, i.e., about 9.6 full-frame passes. The condition tensor would then have near-100% nonzero trajectory coverage, while DragNUWA and MOFA-Video receive only the original 15% sparse points. The paper never reports the nonzero trajectory coverage of the condition tensors, so the 27.82 FVD gap in Table 1 could reflect the amount of motion information provided rather than the superiority of the region-wise representation. Please report coverage statistics and add a controlled comparison in which baselines receive the same padded/dense condition or MotionPro receives only truly sparse unpadded points.
- [§4.4 and Tables 2–3] Hyperparameters are selected on MC-Bench and the final method is then evaluated on the same MC-Bench. Section 4.1 states that the local region size k=8 and the minimal mask ratio rmin=0.95 are 'determined by cross validation,' and Section 4.4 / Figure 6 shows k and rmin being chosen by MD-Vid and Frame Consistency on MC-Bench. The same MC-Bench is subsequently used for the final comparisons in Table 2 and Table 3. This is a selection-on-the-test-set protocol that inflates the reported MD-Img/MD-Vid numbers. Please hold out a portion of MC-Bench for validation and report final results on the held-out portion, or otherwise document that the reported numbers are on a split that was not used for hyperparameter selection.
- [§3.2, Eq. (7), and §4.4] There is a serious train-inference distribution shift that is not addressed. In training, the mask ratio rm is uniformly sampled from [0.95, 1.0], meaning that 95–100% of the k×k regions are active and the condition is an almost-dense flow field. At inference (Section 3.4), users provide one or two sparse trajectories, which are padded to k×k regions. The paper argues in Section 4.4 that 'using small value of rmin will sample more trajectories,' but according to Eq. (7) rm is the probability with which Msel entries are 1, so a smaller rmin reduces the expected number of selected regions. The sentence appears to invert the definition. Please clarify the semantics, justify why training on nearly dense conditions transfers to sparse user inputs, and report experiments with training-time sparsity levels that actually match the inference protocol (e.g., rmin values corresponding to a few selected regions).
- [Tables 1–3 and Supplementary Table 4] All quantitative results are single-run point estimates with no error bars, multiple seeds, or statistical significance tests. This is particularly important for the MC-Bench numbers, which are computed on only 1.1K image-trajectory pairs, and for the human evaluation in Supplementary Table 4, where preference ratios are reported without confidence intervals or inter-annotator agreement. Please provide variance estimates (at least 3 seeds) for the main comparisons and significance tests for the human preference results.
- [§4.1, MD-Vid definition] The MD-Vid metric may be biased toward models trained on DOT-style flows. Training conditions are generated by DOT, and MD-Vid uses CoTracker for trajectory matching; both are dense tracking models from the same family. To rule out that MotionPro's superior MD-Vid scores reflect a favorable interaction between the training tracker and the evaluation tracker, please report trajectory-alignment results using an independent correspondence method (e.g., DIFT-based MD-Img is already reported, but also consider a different video tracker such as RAFT or PointOdyssey) and discuss any differences.
minor comments (6)
- [Throughout] The method name 'DragNUW A' should be written as 'DragNUWA' without a space.
- [§4.4] The explanation of rmin appears reversed relative to Eq. (7); please correct the sentence so that the effect of rmin on the number of selected regions is described consistently with the definition.
- [References] Reference [40] (ZoeDepth) has a malformed author list: 'Diana Wofk Peter Wonka Matthias Müller Shariq Farooq Bhat, Reiner Birkl' should be formatted as separate authors.
- [Figures 7 and 10] Figure captions mention 'Viewed with Acrobat Reader' for animated videos, but the PDF appears to contain only still frames; please point readers to the project page or provide the videos as supplementary material.
- [§3.2 and §3.4] The paper describes the motion mask as derived from optical flow in training but as user-brushed at inference; please state this asymmetry explicitly in Section 3.2 to avoid confusion about the training objective.
- [§6.2] The paper mentions 'in-the-wild' training data and object/camera disentanglement, but does not quantify the proportion of camera-motion vs object-motion samples in WebVid-10M; a brief analysis would help contextualize the disentanglement claim.
Circularity Check
No significant circularity: MotionPro's control conditions are extracted from input video, and no fitted parameter is renamed as a prediction.
full rationale
The paper's derivation chain is self-contained rather than circular. The region-wise trajectory and motion mask are generated from the input video via the off-the-shelf DOT tracker (Eqs. 4-7 and the thresholding paragraph), and the denoising loss in Eq. 3 is the standard score-matching objective with the new conditions injected through modulation and LoRA. There is no equation in which a predicted quantity is defined in terms of itself, nor a fitted parameter that is later reported as a prediction; the choices of k=8 and rmin=0.95 are determined by validation and reported as ablations, which is an evaluation-leakage concern on MC-Bench rather than a circular reduction. The MD-Vid metric uses CoTracker while training conditions come from DOT, and even though both are tracking models, no calculation makes the metric algebraically depend on the training tracker. Self-citations appear only in related-work enumerations and are not load-bearing for the central claim. The WebVid-10M FVD comparison may be affected by uneven condition density (15% sampled trajectories padded to 8x8 blocks for MotionPro versus sparse points for baselines), but that is a benchmark-fairness confound, not a circularity that equates the result to its input by construction.
Assumptions & free parameters
free parameters (5)
- local region size k =
8
- minimal mask ratio rmin =
0.95
- motion mask flow threshold =
average flow magnitude > 1
- LoRA rank =
32
- random mask ratio range =
[0.95, 1.0]
assumptions (5)
- domain assumption Stable Video Diffusion provides a usable pretrained video prior that can be adapted for motion control.
- domain assumption The DOT tracking model accurately estimates optical flow and visibility masks on WebVid-10M training videos.
- domain assumption CoTracker and DIFT provide reliable trajectory references for evaluating motion-trajectory alignment.
- domain assumption MC-Bench annotations represent realistic user intent for motion control.
- domain assumption WebVid-10M videos with optical-flow-derived trajectories are a suitable training distribution for interactive motion control.
Cite this review
Pith. "Pith review of MotionPro: A Precise Motion Controller for Image-to-Video Generation." pith.science (2026). https://pith.science/paper/XT3AL5A2
@misc{pith2026250520287,
author = {Pith},
title = {Pith review of: MotionPro: A Precise Motion Controller for Image-to-Video Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XT3AL5A2}},
note = {Machine review of arXiv:2505.20287}
}
read the original abstract
Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining movement region, leading to coarse motion control and failing to disentangle object and camera moving. To alleviate these, we present MotionPro, a precise motion controller that novelly leverages region-wise trajectory and motion mask to regulate fine-grained motion synthesis and identify target motion category (i.e., object or camera moving), respectively. Technically, MotionPro first estimates the flow maps on each training video via a tracking model, and then samples the region-wise trajectories to simulate inference scenario. Instead of extending flow through large Gaussian kernels, our region-wise trajectory approach enables more precise control by directly utilizing trajectories within local regions, thereby effectively characterizing fine-grained movements. A motion mask is simultaneously derived from the predicted flow maps to capture the holistic motion dynamics of the movement regions. To pursue natural motion control, MotionPro further strengthens video denoising by incorporating both region-wise trajectories and motion mask through feature modulation. More remarkably, we meticulously construct a benchmark, i.e., MC-Bench, with 1.1K user-annotated image-trajectory pairs, for the evaluation of both fine-grained and object-level I2V motion control. Extensive experiments conducted on WebVid-10M and MC-Bench demonstrate the effectiveness of MotionPro. Please refer to our project page for more results: https://zhw-zhang.github.io/MotionPro-page/.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Max Bain, Arsha Nagrani, G ¨ul Varol, and Andrew Zisser- man. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. In ICCV, 2021. 5, 12
work page 2021
-
[2]
Vidu: a Highly Consistent, Dy- namic and Skilled Text-to-Video Generator with Diffusion Models
Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a Highly Consistent, Dy- namic and Skilled Text-to-Video Generator with Diffusion Models. arXiv preprint arXiv:2405.04233, 2024. 2
arXiv 2024
-
[3]
Lumiere: A Space-Time Diffusion Model for Video Generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, and Inbar Mosseri. Lumiere: A Space-Time Diffusion Model for Video Generation. arXiv preprint arXiv:2401.12945, 2024. 2
arXiv 2024
-
[4]
Improving Image Generation with Better Captions, 2023
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, and Yunxin Jiao. Improving Image Generation with Better Captions, 2023. 12
work page 2023
-
[5]
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv preprint arXiv:2311.15127, 2023. 2, 3, 5
arXiv 2023
-
[6]
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models. In CVPR, 2023. 2
work page 2023
-
[7]
Video Generation Models as World Simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video Generation Models as World Simulators
-
[8]
Stable- Video: Text-driven Consistency-aware Diffusion Video Edit- ing
Wenhao Chai, Xun Guo, Gaoang Wang, and Yan Lu. Stable- Video: Text-driven Consistency-aware Diffusion Video Edit- ing. In ICCV, 2023. 2
work page 2023
Show all 64 references
-
[9]
Ouroboros-Diffusion: Ex- ploring Consistent Content Generation in Tuning-free Long Video Diffusion
Jingyuan Chen, Fuchen Long, Jie An, Zhaofan Qiu, Ting Yao, Jiebo Luo, and Tao Mei. Ouroboros-Diffusion: Ex- ploring Consistent Content Generation in Tuning-free Long Video Diffusion. In AAAI, 2025. 1
2025
-
[10]
Control- A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning.arXiv preprint arXiv:2305.13840, 2023
Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Ji- ashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. Control- A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning.arXiv preprint arXiv:2305.13840, 2023. 2
2023 arXiv
-
[11]
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction. In ICCV,
-
[12]
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
Zhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao, Wen- gang Zhou, Jiebo Luo, and Tao Mei. Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution. In CVPR, 2024. 1
2024
-
[13]
Structure and Content-Guided Video Synthesis with Diffusion Mod- els
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and Content-Guided Video Synthesis with Diffusion Mod- els. In ICCV, 2023. 1, 2
2023
-
[14]
Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming- Yu Liu, and Yogesh Balaji. Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models. In ICCV, 2023. 2
2023
-
[15]
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Du- val, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. In ECCV, 2024. 2
2024
-
[16]
I2V- Adapter: A General Image-to-Video Adapter for Diffusion Models
Xun Guo, Mingwu Zheng, Liang Hou, Yuan Gao, Yufan Deng, Pengfei Wan, Di Zhang, Yufan Liu, Weiming Hu, Zhengjun Zha, Haibin Huang, and Chongyang Ma. I2V- Adapter: A General Image-to-Video Adapter for Diffusion Models. arXiv preprint arXiv:2312.16693, 2023. 1
2023 arXiv
-
[17]
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Spe- cific Tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Spe- cific Tuning. In ICLR, 2024. 2
2024
-
[18]
Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and Jos ´e Lezama. Photorealistic Video Generation with Diffusion Models. arXiv preprint arXiv:2312.06662, 2023. 1, 2
2023 arXiv
-
[19]
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In NeuIPS, 2017. 5
2017
-
[20]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Sal- imans. Imagen Video: High Definition Video Generation with Diffusion Models. In CVPR, 2022. 1, 2
2022
-
[21]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video Dif- fusion Models. In NeurIPS, 2022
2022
-
[22]
CogVideo: Large-scale Pretraining for Text-to- Video Generation via Transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. CogVideo: Large-scale Pretraining for Text-to- Video Generation via Transformers. In ICLR, 2023. 1, 2
2023
-
[23]
LoRA: Low-Rank Adaptation of Large Language Models
Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR, 2022. 3, 4
2022
-
[24]
PEEKABOO: Interactive Video Generation via Masked- Diffusion
Yash Jain, Anshul Nasery, Vibhav Vineet, and Harkirat Behl. PEEKABOO: Interactive Video Generation via Masked- Diffusion. In CVPR, 2024. 2
2024
-
[25]
Co- Tracker: It is Better to Track Together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- Tracker: It is Better to Track Together. In ECCV, 2024. 5
2024
-
[26]
Elucidating the Design Space of Diffusion-Based Generative Models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the Design Space of Diffusion-Based Generative Models. In NeurIPS, 2022. 3
2022
-
[27]
Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators
Levon Khachatryan, Andranik Movsisyan, Vahram Tade- vosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2Video-Zero: Text- to-Image Diffusion Models are Zero-Shot Video Generators. In ICCV, 2023. 1, 2
2023
-
[28]
AnimateAnywhere: Context-Controllable Hu- man Video Generation with ID-Consistent One-shot Learn- ing
Hengyuan Liu, Xiaodong Chen, Xinchen Liu, Xiaoyan Gu, and Wu Liu. AnimateAnywhere: Context-Controllable Hu- man Video Generation with ID-Consistent One-shot Learn- ing. In HCMA, 2024. 1
2024
-
[29]
VideoStudio: Generating Consistent-Content and Multi- Scene Videos
Fuchen Long, Zhaofan Qiu, Ting Yao, and Tao Mei. VideoStudio: Generating Consistent-Content and Multi- Scene Videos. In ECCV, 2024. 1
2024
-
[30]
VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tie- niu Tan. VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation. In CVPR, 2023. 2
2023
-
[31]
Lewis, and W
Wan-Duo Kurt Ma, J.P. Lewis, and W. Bastiaan Kleijn. Trail- Blazer: Trajectory Control for Diffusion-Based Video Gen- eration. arXiv preprint arXiv:2401.00896, 2023. 2
2023 arXiv
-
[32]
Latte: Latent Diffusion Transformer for Video Generation
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Zi- wei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent Diffusion Transformer for Video Generation. arXiv preprint arXiv:2401.03048, 2024. 2
2024 arXiv
-
[33]
Dense Optical Tracking: Connecting the Dots
Guillaume Le Moing, Jean Ponce, and Cordelia Schmid. Dense Optical Tracking: Connecting the Dots. In CVPR,
-
[34]
ReVideo: Remake a Video with Motion and Content Control
Chong Mou, Mingdeng Cao, Xintao Wang, Zhaoyang Zhang, Ying Shan, and Jian Zhang. ReVideo: Remake a Video with Motion and Content Control. In NeurIPS, 2024. 2
2024
-
[35]
MOFA-Video: Control- lable Image Animation via Generative Motion Field Adap- tions in Frozen Image-to-Video Diffusion Model
Muyao Niu, Xiaodong Cun, Xintao Wang, Yong Zhang, Ying Shan, and Yinqiang Zheng. MOFA-Video: Control- lable Image Animation via Generative Motion Field Adap- tions in Frozen Image-to-Video Diffusion Model. In ECCV,
-
[36]
FateZero: Fus- ing Attentions for Zero-Shot Text-Based Video Editing
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. FateZero: Fus- ing Attentions for Zero-Shot Text-Based Video Editing. In ICCV, 2023. 5
2023
-
[37]
Boosting Diffusion Models with Moving Average Sampling in Frequency Domain
Yurui Qian, Qi Cai, Yingwei Pan, Yehao Li, Ting Yao, Qibin Sun, and Tao Mei. Boosting Diffusion Models with Moving Average Sampling in Frequency Domain. In CVPR, 2024. 1
2024
-
[38]
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
Mengxue Qu, Xiaodong Chen, Wu Liu, Alicia Li, and Yao Zhao. ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models. In CVPR, 2024. 1
2024
-
[39]
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In ICML,
-
[40]
ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
Diana Wofk Peter Wonka Matthias M ¨uller Shariq Fa- rooq Bhat, Reiner Birkl. ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth. In CVPR, 2023. 13
2023
-
[41]
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling. In ACM SIG- GRAPH, 2024. 2
2024
-
[42]
Yujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan, Han- shu Yan, Wenqing Zhang, Vincent Y . F. Tan, and Song Bai. DragDiffusion: Harnessing Diffusion Models for Interactive Point-Based Image Editing. In CVPR, 2024. 5, 12, 13
2024
-
[43]
Make-a-Video: Text-to-Video Generation without Text-Video Data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taig- man. Make-a-Video: Text-to-Video Generation without Text-Video Data. In ICLR, 2023. 1, 2
2023
-
[44]
Emergent Correspondence from Image Diffusion
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent Correspondence from Image Diffusion. In NeurIPS, 2023. 5
2023
-
[45]
FVD: A new Metric for Video Generation
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. FVD: A new Metric for Video Generation. In ICLR Deep- GenStruct Workshop, 2019. 5
2019
-
[46]
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
Siqi Wan, Jingwen Chen, Yingwei Pan, Ting Yao, and Tao Mei. Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On. In ICLR, 2025. 1
2025
-
[47]
ModelScope Text-to-Video Technical Report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xi- ang Wang, and Shiwei Zhang. ModelScope Text-to-Video Technical Report. arXiv preprint arXiv:2308.06571 , 2023. 1, 2
2023 arXiv
-
[48]
Boximator: Gen- erating Rich and Controllable Motions for Video Synthesis
Jiawei Wang, Yuchen Zhang, Jiaxin Zou, Yan Zeng, Guo- qiang Wei, Liping Yuan, and Hang Li. Boximator: Gen- erating Rich and Controllable Motions for Video Synthesis. arXiv preprint arXiv:2402.01566, 2024. 2
2024 arXiv
-
[49]
VideoComposer: Compositional Video Synthesis with Motion Controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Ji- uniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jin- gren Zhou. VideoComposer: Compositional Video Synthesis with Motion Controllability. In NeurIPS, 2023. 1, 2
2023
-
[50]
MotionCtrl: A Uni- fied and Flexible Motion Controller for Video Generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. MotionCtrl: A Uni- fied and Flexible Motion Controller for Video Generation. In ACM SIGGRAPH, 2024. 2
2024
-
[51]
Mo- tionBooth: Motion-Aware Customized Text-to-Video Gener- ation
Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou, Yining Li, Yunhai Tong, and Kai Chen. Mo- tionBooth: Motion-Aware Customized Text-to-Video Gener- ation. In NeurIPS, 2024. 2
2024
-
[52]
DragAnything: Motion Control for Any- thing using Entity Representation
Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. DragAnything: Motion Control for Any- thing using Entity Representation. In ECCV, 2024. 1, 2, 6, 12, 13
2024
-
[53]
DynamiCrafter: Animating Open-domain Images with Video Diffusion Pri- ors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Xin- tao Wang, Tien-Tsin Wong, and Ying Shan. DynamiCrafter: Animating Open-domain Images with Video Diffusion Pri- ors. In ECCV, 2024. 2
2024
-
[54]
Hi3D: Pursuing High- Resolution Image-to-3D Generation with Video Diffusion Models
Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao, Zhineng Chen, Chong-Wah Ngo, and Tao Mei. Hi3D: Pursuing High- Resolution Image-to-3D Generation with Video Diffusion Models. In ACM MM, 2024. 1
2024
-
[55]
Direct-a-Video: Customized Video Generation with User- Directed Camera Movement and Object Motion
Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao. Direct-a-Video: Customized Video Generation with User- Directed Camera Movement and Object Motion. In ACM SIGGRAPH, 2024. 2
2024
-
[56]
CogVideoX: Text-to-Video Diffu- sion Models with An Expert Transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xi- aohan Zhang, Guanyu Feng, Da Yin, Xiaotao Gu, Yuxuan Zhang, Weihan Wang, Yean Cheng, Ting Liu, Bin Xu, Yux- iao Dong, and Jie Tang. CogVideoX: Text-to-Video Diffu- sion M...
2024 arXiv
-
[57]
DragNUW A: Fine-Grained Control in Video Generation by Integrating Text, Image, and Trajectory
Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan. DragNUW A: Fine-Grained Control in Video Generation by Integrating Text, Image, and Trajectory. arXiv preprint arXiv:2308.08089, 2023. 1, 2, 5, 12
2023 arXiv
-
[58]
Make Pixels Dance: High- Dynamic Video Generation
Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. Make Pixels Dance: High- Dynamic Video Generation. In CVPR, 2024. 2
2024
-
[59]
Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding Conditional Control to Text-to-Image Diffusion Models. In ICCV, 2023. 4
2023
-
[60]
I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou. I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models. arXiv preprint arXiv:2311.04145, 2023. 2
2023 arXiv
-
[61]
ControlVideo: Training-Free Controllable Text-to-Video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. ControlVideo: Training-Free Controllable Text-to-Video Generation. In ICLR, 2024. 1, 2
2024
-
[62]
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffu- sion Models
Zhongwei Zhang, Fuchen Long, Yingwei Pan, Zhaofan Qiu, Ting Yao, Yang Cao, and Tao Mei. TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffu- sion Models. In CVPR, 2024. 1
2024
-
[63]
Supplementary Material The supplementary material contains: 1) the dataset details of MC-Bench; 2) baseline choices and experimental details
-
[64]
in- the-wild
the human evaluation of motion control; 4) robustness of motion mask; 5) the application of camera control; 6) runtime comparison; 7) ablation on control signals. 6.1. Dataset Details of MC-Bench The proposed MC-Bench consists of412 high-quality refer- ence images and correspo...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.