REVIEW 5 major objections 5 minor 75 references
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read ChoreoMuse is a two-stage diffusion framework that generates style-controlled dance videos from any music and reference image, using SMPL body parameters as the bridge between music and pixels.
desk verdict A well-built music-to-dance video system, but its style-adherence SOTA claim rests on unvalidated metrics and thin baselines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the SMPL parametric body model, a low-dimensional representation of human shape and pose, used in two parameterizations: the original compact form for the video stage and a 6D-rotation form with foot-contact labels for the choreography stage. Stage one uses a DDPM conditioned on a fused embedding made from MotionTune's contrastively trained audio-motion representation, Jukebox features, and a style embedding produced by a music classifier and text encoder. Stage two renders the transferred SMPL mesh into depth, normal, segmentation, and foot-contact maps, fuses them with Multi-Layer Motion Fusion, and feeds them through cross-attention to a U-Net video diffusion model, with a silhouette-based shape-alignment step adjusting the SMPL body parameters to match the reference person's contours.
What would settle it
Run a preregistered human study in which independent choreographers pick which generated clip best matches a target choreography style, and check whether the CSAS and MSAS rankings agree with their picks when the feature extractor in CSAS is replaced by a frozen, independently trained motion encoder; if the rankings flip, the style-alignment claim is an artifact of the metric.
Extended reading notes
Core claim
The central claim is that decomposing music-to-dance video generation into two diffusion stages connected by an explicit SMPL-format 3D dance sequence removes the resolution and background constraints that plague direct music-to-video methods while preserving beat adherence and enabling style transfer. In the first stage, a denoising diffusion model generates a 3D dance sequence in a 6D-rotation variant of SMPL, conditioned on music embeddings and a classifier-driven choreography style; in the second, a video diffusion model animates the reference person from that sequence rendered into depth, normal, segmentation, and foot-contact maps. The authors report that this design outperforms the prior direct music-to-dance baseline and several pose-guided animation methods on video quality, and outperforms motion-generation baselines on beat alignment, dance diversity, and the two new style-alignment scores.
Load-bearing premise
The style-adherence claim rests on two new scores defined by the same authors: one uses a style classifier trained on the same AIST++ dataset used to train the model, and the other uses a feature extractor that is never specified plus a hand-chosen decay parameter, so the whole style-match result depends on these scores tracking what humans actually perceive as style.
Editorial extensions
If this is right
- Music-to-dance video generation becomes resolution-independent: the output video can match the resolution and environment of any reference image rather than a fixed low-resolution canvas.
- Users can choreograph the same music in different styles, and can reuse one choreography with different reference people or backgrounds, because the 3D dance sequence and the final video are generated in separate stages.
- The contrastively trained MotionTune encoder should improve beat alignment and motion diversity relative to conditioning on generic audio features alone.
- The MSAS and CSAS scores give later work a quantitative way to compare style adherence, extending evaluation beyond beat scores and physical-plausibility checks.
Reading between the lines
- Because the SMPL intermediate decouples motion from appearance, the same trained modules could plausibly be retargeted to stylized avatars or non-human subjects, which the paper's toy, comic-character, and oil-painting examples already hint at.
- A natural extension would let users supply reference dance clips instead of choosing a classifier-defined genre, so the style controller could imitate an arbitrary choreographic vocabulary rather than only the styles present in AIST++.
- The silhouette-based shape alignment suggests a testable recipe: fit SMPL to a reference photo once, then reuse those fitted parameters across many music inputs, isolating choreography quality from identity-preservation quality in evaluation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. ChoreoMuse is a two-stage diffusion framework for music-to-dance video generation. The first stage generates 3D dance sequences in an SMPL-based 6D rotation representation from audio, using a contrastively trained music encoder (MotionTune), a music-style classifier, and a choreography style controller. The second stage renders a reference person into a video guided by the generated 3D sequence, using silhouette/keypoint shape alignment and multi-layer motion fusion. The paper introduces two style metrics (MSAS and CSAS) and reports experiments on AIST++ and TikTok datasets, claiming state-of-the-art video quality, beat alignment, diversity, and style adherence.
Significance. The system is a plausible engineering contribution: using SMPL parameters as an intermediate between music and video is a reasonable design choice that bypasses resolution limits, and MotionTune is a sensible mechanism for aligning audio and motion embeddings. If the style-control and quality claims were rigorously validated, the work would be useful to the computational choreography and human animation communities. However, the paper's headline claims currently outrun its evidence: the style metrics are not validated, the only music-conditioned video baseline is the authors' own DabFusion, and the TikTok comparisons are against pose-guided methods on a different task. The technical scaffolding is interesting, but the evaluation needs substantial strengthening before the state-of-the-art claim can be accepted.
major comments (5)
- [Sec. 4.3, Eq. (18), Table 2] The claim of state-of-the-art style adherence is unsupported. CSAS is reported only for ChoreoMuse in Table 2 (all baselines have '-'), so there is no comparison on this dimension. Moreover, Eq. (18) defines CSAS in terms of an unspecified feature embedding phi(x) and an unreported decay parameter alpha; without these choices and a sensitivity analysis, the metric is not reproducible and its value of 0.84 cannot be interpreted. The paper must specify phi and alpha, and either provide a baseline comparison on CSAS or restrict the claim to internal evaluation.
- [Sec. 4.1, Sec. 4.3, Eq. (15)] MSAS is likely inflated by distributional overlap rather than style fidelity. The multi-class style classifier used in MSAS is trained on AIST++ (Sec. 4.1), the same dataset used to train the dance generator, so generated sequences are in-distribution for the classifier by construction. High MSAS may simply reflect that the generated dances resemble the training set, not that they match the intended musical style. To support the style-adherence claim, the authors should correlate MSAS with human perceptual judgments or evaluate with an independently trained style classifier on out-of-distribution data.
- [Sec. 4.2, Table 1] The video-quality comparison on TikTok is not an apples-to-apples evaluation of music-to-dance generation. DisCo, MagicAnimate, and Animate Anyone are pose-guided image animation baselines: they are conditioned on ground-truth poses rather than on music, so the task is different from ChoreoMuse's music-conditioned video generation. The reported margins on TikTok (e.g., PSNR 29.85 vs 29.49, FVD 165.4 vs 173.5) therefore do not establish superiority for the music-to-dance task. On AIST++, where the task matches, the only baseline is DabFusion, the authors' own prior system; this single comparison does not justify the phrase 'state-of-the-art across multiple dimensions' in the abstract.
- [Sec. 4.4, Human Evaluations] The user study does not validate the proposed metrics as claimed. The text says the study was conducted 'to validate these metrics', but it only reports whether participants judged generated videos as matching music style (83.2%) and choreography style (76.8%). It never correlates these human judgments with MSAS or CSAS scores, so it cannot establish that either metric tracks human perception. In addition, the 80% success threshold is ad hoc and no confidence intervals or significance tests are supplied. The authors should either report a correlation analysis between human ratings and the proposed metrics or remove the validation claim.
- [Tables 1–5] All quantitative results are reported without variance or significance testing. For example, Table 1 reports FVD values 176.3 (MagicAnimate) vs 165.4 (Ours) on TikTok, and Table 2 reports BAS 0.26 (EDGE) vs 0.28 (Ours), but without multiple runs, confidence intervals, or paired significance tests these differences may not be reliable. Given that several comparisons involve small margins, the paper should include error bars and statistical tests for the headline metrics (PSNR, SSIM, LPIPS, FVD, PFC, BAS, diversity, MSAS, CSAS).
minor comments (5)
- [Sec. 3.3, Eq. (11)] The symbol 'B' in the equation 'x_{t-1} B m⊙q(...) + ...' appears to be a typesetting error; it should probably be an equality or assignment symbol.
- [Sec. 3.3] The paper mentions that the model can generate a 7.5-second clip by constraining the first 2.5 seconds of the new sequence, but all training clips are 5 seconds; a brief description of how the model handles variable-length inference would help reproducibility.
- [Sec. 4.4] The sentence 'we propose our CSAS metric as a baseline for future style-controllable research' is confusing because a baseline is a comparison method, not a metric; rephrase to describe CSAS as a proposed evaluation protocol.
- [Sec. 4.3] The description of the MSAS classifier and the CSAS centroid computation would benefit from stating the dimensionality of the embeddings and the number of style classes, since these details affect the interpretation of the scores.
- [Sec. 4.4] The user study section states that 'Additional user study details are provided in the supplementary materials', but no supplement is included with the manuscript; if this is a journal submission, the supplementary material should be provided.
Circularity Check
No significant circularity: core results rest on external metrics; the author-designed style metrics and self-cited DabFusion baseline are validity/positioning concerns, not formal reductions.
full rationale
The paper's main generation pipeline is evaluated with standard, externally defined metrics: PSNR, SSIM, LPIPS, FVD (Table 1), PFC, Dist_k, Dist_g, BAS (Table 2), with ablations in Tables 3-5. These do not reduce to the model's training inputs by construction. The two new metrics MSAS (Eqs. 15-16) and CSAS (Eq. 18) are author-introduced; MSAS uses a style classifier trained on AIST++, the same dataset used to train the generator, and CSAS depends on an unspecified embedding phi and an unreported decay alpha. These are legitimate concerns about metric validity and reproducibility, and the human study in Sec. 4.4 does not correlate participant judgments with MSAS/CSAS scores. However, no equation in the paper defines the metric output as equivalent to a fitted parameter or model output by construction; the concern is that the metrics may not measure what they claim, not that the claims are circular. The only self-citation, DabFusion [56], is used as a baseline and to motivate the task; since ChoreoMuse's quantitative comparisons and ablations are independent of that citation, it is not load-bearing. Overall, no specific derivation chain reduces to its own inputs.
Assumptions & free parameters
free parameters (4)
- alpha (CSAS decay parameter) =
not reported
- lambda_pos, lambda_vel, lambda_foot =
not reported
- lambda_kpt, lambda_sil =
not reported
- 80% success threshold in user evaluation =
0.80
assumptions (5)
- standard math DDPM forward and reverse processes are correctly implemented as in Ho et al.
- domain assumption AIST++ music and motion pairs provide reliable genre and choreography style labels for contrastive training and classification.
- domain assumption SMPL and 6D rotation pose parameters carry enough information to drive photorealistic video generation.
- domain assumption A music classifier can reliably predict the genre, and the mapped choreography style is acceptable.
- domain assumption The Champ and TikTok video datasets are representative enough to train a video generator that handles any reference individual at any resolution.
Cite this review
Pith. "Pith review of ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion." pith.science (2026). https://pith.science/paper/4IP7SKWF
@misc{pith2026250719836,
author = {Pith},
title = {Pith review of: ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion},
year = {2026},
howpublished = {\url{https://pith.science/paper/4IP7SKWF}},
note = {Machine review of arXiv:2507.19836}
}
read the original abstract
Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance videos that harmonize with both musical rhythm and user-defined choreography styles, limiting their applicability in real-world creative contexts. To address this gap, we introduce ChoreoMuse, a diffusion-based framework that uses SMPL format parameters and their variation version as intermediaries between music and video generation, thereby overcoming the usual constraints imposed by video resolution. Critically, ChoreoMuse supports style-controllable, high-fidelity dance video generation across diverse musical genres and individual dancer characteristics, including the flexibility to handle any reference individual at any resolution. Our method employs a novel music encoder MotionTune to capture motion cues from audio, ensuring that the generated choreography closely follows the beat and expressive qualities of the input music. To quantitatively evaluate how well the generated dances match both musical and choreographic styles, we introduce two new metrics that measure alignment with the intended stylistic cues. Extensive experiments confirm that ChoreoMuse achieves state-of-the-art performance across multiple dimensions, including video quality, beat alignment, dance diversity, and style adherence, demonstrating its potential as a robust solution for a wide range of creative applications. Video results can be found on our project page: https://choreomuse.github.io.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Emre Aksan, Manuel Kaufmann, and Otmar Hilliges. 2019. Structured prediction helps 3d human motion modelling. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7144–7153
work page 2019
-
[2]
Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan. 2023. Person image synthesis via denoising diffusion model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5968–5976
work page 2023
-
[3]
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. 2023. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22563–22575
work page 2023
-
[4]
Judith Butepage, Michael J Black, Danica Kragic, and Hedvig Kjellstrom. 2017. Deep representation learning for human motion prediction and classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6158–6166
work page 2017
-
[5]
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2017. Realtime multi- person 2d pose estimation using part affinity fields. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7291–7299
work page 2017
-
[6]
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022. HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection. In ICASSP 2022-2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing . IEEE, 646–650
work page 2022
-
[7]
Yimian Dai, Fabian Gieseke, Stefan Oehmcke, Yiquan Wu, and Kobus Barnard
-
[8]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186
2019
Show all 75 references
-
[9]
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341 (2020)
2020 arXiv
-
[10]
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang
-
[11]
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. 2023. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7346–7356
2023
-
[12]
Joao P Ferreira, Thiago M Coutinho, Thiago L Gomes, José F Neto, Rafael Azevedo, Renato Martins, and Erickson R Nascimento. 2021. Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio. Computers & Graphics 94 (2021), 11–21
2021
-
[13]
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. 2019. Learning individual styles of conversational gesture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3497–3506
2019
-
[14]
Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, and Xinchao Wang. 2023. Tm2d: Bimodality driven 3d dance generation via music-text integration. InProceedings of the IEEE/CVF International Conference on Computer Vision . 9942–9952
2023
-
[15]
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng
-
[16]
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. 2023. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725 (2023)
2023 arXiv
-
[17]
William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, and Frank Wood. 2022. Flexible diffusion modeling of long videos.Advances in Neural Information Processing Systems 35 (2022), 27953–27965
2022
-
[18]
Alejandro Hernandez, Jurgen Gall, and Francesc Moreno-Noguer. 2019. Human motion prediction via spatio-temporal inpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7134–7143
2019
-
[19]
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303 (2022)
2022 arXiv
-
[20]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020), 6840–6851
2020
-
[21]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video diffusion models. Advances in Neu- ral Information Processing Systems 35 (2022), 8633–8646
2022
-
[22]
Daniel Holden, Jun Saito, and Taku Komura. 2016. A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics 35, 4 (2016), 1–11
2016
-
[23]
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2022. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868 (2022)
2022 arXiv
-
[24]
Alain Hore and Djemel Ziou. 2010. Image quality metrics: PSNR vs. SSIM. In 2010 20th International Conference on Pattern Recognition . IEEE, 2366–2369
2010
-
[25]
Li Hu. 2024. Animate anyone: Consistent and controllable image-to-video syn- thesis for character animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8153–8163
2024
-
[26]
Ruozi Huang, Huang Hu, Wei Wu, Kei Sawada, Mi Zhang, and Daxin Jiang
-
[27]
Yasamin Jafarian and Hyun Soo Park. 2021. Learning high fidelity depths of dressed humans by watching social media dance videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12753–12762
2021
-
[28]
Hsuan-Kai Kao and Li Su. 2020. Temporally guided music-to-body-movement generation. In Proceedings of the 28th ACM International Conference on Multimedia. 147–155
2020
-
[29]
Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher- Shlizerman. 2023. Dreampose: Fashion video synthesis with stable diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 22680– 22690
2023
-
[30]
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Hen- schel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. 2023. Text2video-zero: Text-to-image diffusion models are zero-shot video genera- tors. In Proceedings of the IEEE/CVF International Conference on...
2023
-
[31]
Jinwoo Kim, Heeseok Oh, Seongjean Kim, Hoseok Tong, and Sanghoon Lee. 2022. A brand new dance partner: Music-conditioned pluralistic dancing controlled by multiple dance genres. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3490–3500
2022
-
[32]
Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)
2013 arXiv
-
[33]
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley. 2020. Panns: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020), 2880–2894
2020
-
[34]
Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. 2019. Dancing to music. Advances in Neural Information Processing Systems 32 (2019)
2019
-
[35]
Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. 2021. Ai choreog- rapher: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 13401–13412
2021
-
[36]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015. SMPL: A Skinned Multi-Person Linear Model. ACM Transactions on Graphics 34, 6 (2015)
2015
-
[37]
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion proba- bilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11461–11471
2022
-
[38]
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. 2024. Follow your pose: Pose-guided text-to-video generation using pose-free videos. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 4117–4125
2024
-
[39]
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python.. In SciPy. 18–24
2015
-
[40]
Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min. 2023. Conditional image-to-video generation with latent flow diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18444–18455
2023
-
[41]
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
2018
-
[42]
Mathis Petrovich, Michael J Black, and Gül Varol. 2021. Action-conditioned 3D human motion synthesis with transformer VAE. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10985–10995
2021
-
[43]
Mathis Petrovich, Michael J Black, and Gül Varol. 2022. Temos: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision. Springer, 480–497
2022
-
[44]
Qiaosong Qi, Le Zhuo, Aixi Zhang, Yue Liao, Fei Fang, Si Liu, and Shuicheng Yan
-
[45]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learni...
2021
-
[46]
Xuanchi Ren, Haoran Li, Zijian Huang, and Qifeng Chen. 2020. Self-supervised dance video synthesis conditioned on music. In Proceedings of the 28th ACM International Conference on Multimedia . 46–54
2020
-
[47]
Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman
-
[48]
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. 2022. Bailando: 3d dance generation by actor- critic gpt with choreographic memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11050–11059
2022
-
[49]
In Proceedings of the 31st ACM International Conference on Multimedia
Diffdance: Cascaded human motion diffusion model for dance generation. In Proceedings of the 31st ACM International Conference on Multimedia. 1374–1382
-
[50]
Guofei Sun, Yongkang Wong, Zhiyong Cheng, Mohan S Kankanhalli, Weidong Geng, and Xiangdong Li. 2020. Deepdance: music-to-dance motion choreography with adversarial learning. IEEE Transactions on Multimedia 23 (2020), 497–509
2020
-
[51]
Taoran Tang, Jia Jia, and Hanyang Mao. 2018. Dance with melody: An lstm- autoencoder approach to music-oriented dance synthesis. In Proceedings of the 26th ACM International Conference on Multimedia . 1598–1606
2018
-
[52]
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or
-
[53]
Jonathan Tseng, Rodrigo Castellon, and Karen Liu. 2023. Edge: Editable dance generation from music. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 448–458
2023
-
[54]
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. 2018. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717 (2018)
2018 arXiv
-
[55]
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli
-
[56]
Xuanchen Wang, Heng Wang, Dongnan Liu, and Weidong Cai. 2025. Dance Any Beat: Blending Beats with Visuals in Dance Video Generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
2025
-
[57]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Transactions on Image Processing 13, 4 (2004), 600–612
2004
-
[58]
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. 2022. Wav2clip: Learning robust audio representations from clip. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 4563–4567
2022
-
[59]
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech an...
2023
-
[60]
In European Conference on Computer Vision
Motionclip: Exposing human motion generation to clip space. In European Conference on Computer Vision . Springer, 358–374
-
[61]
Nelson Yalta, Shinji Watanabe, Kazuhiro Nakadai, and Tetsuya Ogata. 2019. Weakly-supervised deep recurrent neural networks for basic dance step genera- tion. In 2019 International Joint Conference on Neural Networks . IEEE, 1–8
2019
-
[62]
Zijie Ye, Haozhe Wu, Jia Jia, Yaohua Bu, Wei Chen, Fanbo Meng, and Yanfeng Wang. 2020. Choreonet: Towards music to dance synthesis with choreographic action unit. InProceedings of the 28th ACM International Conference on Multimedia. 744–752
2020
-
[63]
Tan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2024. Disco: Disentangled control for realistic human dance generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2024
-
[64]
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5745–5753
2019
-
[65]
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. 2024. Champ: Controllable and consistent human image animation with 3d parametric guidance. In European Conference on Computer Vision. Springer, 145–162
2024
-
[66]
Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang, Ming Shao, and Siyu Xia. 2022. Music2dance: Dancenet for music-driven dance generation. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 2 (2022), 1–21
2022
-
[68]
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2024. Magicanimate: Temporally consistent human image animation using diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2024
-
[71]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[72]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 586–595
-
[2015]
In International Conference on Machine Learning
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265
-
[2018]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Audio to body dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7574–7583
-
[2020]
arXiv preprint arXiv:2006.06119 (2020)
Dance revolution: Long-term dance generation with music via curriculum learning. arXiv preprint arXiv:2006.06119 (2020)
2020 arXiv
-
[2021]
InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
Attentional feature fusion. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 3560–3569
-
[2022]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Generating diverse and natural 3d human motions from text. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5152– 5161
-
[2023]
InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing
Clap learning audio concepts from natural language supervision. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 1–5
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.