Pith. sign in

REVIEW 5 major objections 5 minor 75 references

ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion

T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read ChoreoMuse is a two-stage diffusion framework that generates style-controlled dance videos from any music and reference image, using SMPL body parameters as the bridge between music and pixels.

desk verdict A well-built music-to-dance video system, but its style-adherence SOTA claim rests on unvalidated metrics and thin baselines. read the letter →

arxiv 2507.19836 v1 pith:4IP7SKWF submitted 2025-07-26 cs.GR cs.AIcs.CVcs.MMcs.SD

classification cs.GRcs.AIcs.CVcs.MMcs.SD
keywords music-to-dancegenerationimage-to-videodiffusionmodelSMPLparametricbodybeatalignmentstyletransfercontrastivelearningdancevideosynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ChoreoMuse is a two-stage diffusion framework that generates dance videos from any piece of music and a single reference image, with user-selectable choreography style. The paper's central claim is that 3D body parameters in the SMPL format are a better intermediary between music and pixels than the optical-flow or keypoint representations used by earlier systems, because they are resolution-independent and carry enough structure to guide high-fidelity video. A second claim is that a dedicated music encoder, MotionTune, trained with contrastive audio-motion alignment, captures dance-relevant cues that generic audio features miss, producing stronger beat alignment and greater diversity. The paper also introduces two metrics, MSAS and CSAS, to measure alignment with musical and choreographic style, and reports that ChoreoMuse beats prior music-to-dance video and pose-guided animation methods on all reported benchmarks.

What carries the argument

The load-bearing object is the SMPL parametric body model, a low-dimensional representation of human shape and pose, used in two parameterizations: the original compact form for the video stage and a 6D-rotation form with foot-contact labels for the choreography stage. Stage one uses a DDPM conditioned on a fused embedding made from MotionTune's contrastively trained audio-motion representation, Jukebox features, and a style embedding produced by a music classifier and text encoder. Stage two renders the transferred SMPL mesh into depth, normal, segmentation, and foot-contact maps, fuses them with Multi-Layer Motion Fusion, and feeds them through cross-attention to a U-Net video diffusion model, with a silhouette-based shape-alignment step adjusting the SMPL body parameters to match the reference person's contours.

What would settle it

Run a preregistered human study in which independent choreographers pick which generated clip best matches a target choreography style, and check whether the CSAS and MSAS rankings agree with their picks when the feature extractor in CSAS is replaced by a frozen, independently trained motion encoder; if the rankings flip, the style-alignment claim is an artifact of the metric.

Watch

Extended reading notes

Core claim

The central claim is that decomposing music-to-dance video generation into two diffusion stages connected by an explicit SMPL-format 3D dance sequence removes the resolution and background constraints that plague direct music-to-video methods while preserving beat adherence and enabling style transfer. In the first stage, a denoising diffusion model generates a 3D dance sequence in a 6D-rotation variant of SMPL, conditioned on music embeddings and a classifier-driven choreography style; in the second, a video diffusion model animates the reference person from that sequence rendered into depth, normal, segmentation, and foot-contact maps. The authors report that this design outperforms the prior direct music-to-dance baseline and several pose-guided animation methods on video quality, and outperforms motion-generation baselines on beat alignment, dance diversity, and the two new style-alignment scores.

Load-bearing premise

The style-adherence claim rests on two new scores defined by the same authors: one uses a style classifier trained on the same AIST++ dataset used to train the model, and the other uses a feature extractor that is never specified plus a hand-chosen decay parameter, so the whole style-match result depends on these scores tracking what humans actually perceive as style.

Editorial extensions

If this is right

  • Music-to-dance video generation becomes resolution-independent: the output video can match the resolution and environment of any reference image rather than a fixed low-resolution canvas.
  • Users can choreograph the same music in different styles, and can reuse one choreography with different reference people or backgrounds, because the 3D dance sequence and the final video are generated in separate stages.
  • The contrastively trained MotionTune encoder should improve beat alignment and motion diversity relative to conditioning on generic audio features alone.
  • The MSAS and CSAS scores give later work a quantitative way to compare style adherence, extending evaluation beyond beat scores and physical-plausibility checks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the SMPL intermediate decouples motion from appearance, the same trained modules could plausibly be retargeted to stylized avatars or non-human subjects, which the paper's toy, comic-character, and oil-painting examples already hint at.
  • A natural extension would let users supply reference dance clips instead of choosing a classifier-defined genre, so the style controller could imitate an arbitrary choreographic vocabulary rather than only the styles present in AIST++.
  • The silhouette-based shape alignment suggests a testable recipe: fit SMPL to a reference photo once, then reuse those fitted parameters across many music inputs, isolating choreography quality from identity-preservation quality in evaluation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. ChoreoMuse is a two-stage diffusion framework for music-to-dance video generation. The first stage generates 3D dance sequences in an SMPL-based 6D rotation representation from audio, using a contrastively trained music encoder (MotionTune), a music-style classifier, and a choreography style controller. The second stage renders a reference person into a video guided by the generated 3D sequence, using silhouette/keypoint shape alignment and multi-layer motion fusion. The paper introduces two style metrics (MSAS and CSAS) and reports experiments on AIST++ and TikTok datasets, claiming state-of-the-art video quality, beat alignment, diversity, and style adherence.

Significance. The system is a plausible engineering contribution: using SMPL parameters as an intermediate between music and video is a reasonable design choice that bypasses resolution limits, and MotionTune is a sensible mechanism for aligning audio and motion embeddings. If the style-control and quality claims were rigorously validated, the work would be useful to the computational choreography and human animation communities. However, the paper's headline claims currently outrun its evidence: the style metrics are not validated, the only music-conditioned video baseline is the authors' own DabFusion, and the TikTok comparisons are against pose-guided methods on a different task. The technical scaffolding is interesting, but the evaluation needs substantial strengthening before the state-of-the-art claim can be accepted.

major comments (5)
  1. [Sec. 4.3, Eq. (18), Table 2] The claim of state-of-the-art style adherence is unsupported. CSAS is reported only for ChoreoMuse in Table 2 (all baselines have '-'), so there is no comparison on this dimension. Moreover, Eq. (18) defines CSAS in terms of an unspecified feature embedding phi(x) and an unreported decay parameter alpha; without these choices and a sensitivity analysis, the metric is not reproducible and its value of 0.84 cannot be interpreted. The paper must specify phi and alpha, and either provide a baseline comparison on CSAS or restrict the claim to internal evaluation.
  2. [Sec. 4.1, Sec. 4.3, Eq. (15)] MSAS is likely inflated by distributional overlap rather than style fidelity. The multi-class style classifier used in MSAS is trained on AIST++ (Sec. 4.1), the same dataset used to train the dance generator, so generated sequences are in-distribution for the classifier by construction. High MSAS may simply reflect that the generated dances resemble the training set, not that they match the intended musical style. To support the style-adherence claim, the authors should correlate MSAS with human perceptual judgments or evaluate with an independently trained style classifier on out-of-distribution data.
  3. [Sec. 4.2, Table 1] The video-quality comparison on TikTok is not an apples-to-apples evaluation of music-to-dance generation. DisCo, MagicAnimate, and Animate Anyone are pose-guided image animation baselines: they are conditioned on ground-truth poses rather than on music, so the task is different from ChoreoMuse's music-conditioned video generation. The reported margins on TikTok (e.g., PSNR 29.85 vs 29.49, FVD 165.4 vs 173.5) therefore do not establish superiority for the music-to-dance task. On AIST++, where the task matches, the only baseline is DabFusion, the authors' own prior system; this single comparison does not justify the phrase 'state-of-the-art across multiple dimensions' in the abstract.
  4. [Sec. 4.4, Human Evaluations] The user study does not validate the proposed metrics as claimed. The text says the study was conducted 'to validate these metrics', but it only reports whether participants judged generated videos as matching music style (83.2%) and choreography style (76.8%). It never correlates these human judgments with MSAS or CSAS scores, so it cannot establish that either metric tracks human perception. In addition, the 80% success threshold is ad hoc and no confidence intervals or significance tests are supplied. The authors should either report a correlation analysis between human ratings and the proposed metrics or remove the validation claim.
  5. [Tables 1–5] All quantitative results are reported without variance or significance testing. For example, Table 1 reports FVD values 176.3 (MagicAnimate) vs 165.4 (Ours) on TikTok, and Table 2 reports BAS 0.26 (EDGE) vs 0.28 (Ours), but without multiple runs, confidence intervals, or paired significance tests these differences may not be reliable. Given that several comparisons involve small margins, the paper should include error bars and statistical tests for the headline metrics (PSNR, SSIM, LPIPS, FVD, PFC, BAS, diversity, MSAS, CSAS).
minor comments (5)
  1. [Sec. 3.3, Eq. (11)] The symbol 'B' in the equation 'x_{t-1} B m⊙q(...) + ...' appears to be a typesetting error; it should probably be an equality or assignment symbol.
  2. [Sec. 3.3] The paper mentions that the model can generate a 7.5-second clip by constraining the first 2.5 seconds of the new sequence, but all training clips are 5 seconds; a brief description of how the model handles variable-length inference would help reproducibility.
  3. [Sec. 4.4] The sentence 'we propose our CSAS metric as a baseline for future style-controllable research' is confusing because a baseline is a comparison method, not a metric; rephrase to describe CSAS as a proposed evaluation protocol.
  4. [Sec. 4.3] The description of the MSAS classifier and the CSAS centroid computation would benefit from stating the dimensionality of the embeddings and the number of style classes, since these details affect the interpretation of the scores.
  5. [Sec. 4.4] The user study section states that 'Additional user study details are provided in the supplementary materials', but no supplement is included with the manuscript; if this is a journal submission, the supplementary material should be provided.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: core results rest on external metrics; the author-designed style metrics and self-cited DabFusion baseline are validity/positioning concerns, not formal reductions.

full rationale

The paper's main generation pipeline is evaluated with standard, externally defined metrics: PSNR, SSIM, LPIPS, FVD (Table 1), PFC, Dist_k, Dist_g, BAS (Table 2), with ablations in Tables 3-5. These do not reduce to the model's training inputs by construction. The two new metrics MSAS (Eqs. 15-16) and CSAS (Eq. 18) are author-introduced; MSAS uses a style classifier trained on AIST++, the same dataset used to train the generator, and CSAS depends on an unspecified embedding phi and an unreported decay alpha. These are legitimate concerns about metric validity and reproducibility, and the human study in Sec. 4.4 does not correlate participant judgments with MSAS/CSAS scores. However, no equation in the paper defines the metric output as equivalent to a fitted parameter or model output by construction; the concern is that the metrics may not measure what they claim, not that the claims are circular. The only self-citation, DabFusion [56], is used as a baseline and to motivate the task; since ChoreoMuse's quantitative comparisons and ablations are independent of that citation, it is not load-bearing. Overall, no specific derivation chain reduces to its own inputs.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The ledger lists four hand-chosen parameters that directly affect the reported results but are not given values, plus five background assumptions about datasets and representations. No new physical or conceptual entities are introduced; MotionTune and the new metrics are constructed modules, not postulated entities.

free parameters (4)
  • alpha (CSAS decay parameter) = not reported
    Appears in Eq. (18) as the decay parameter of the Choreography Style Alignment Score; no value or selection method is given, and the reported CSAS values depend on it.
  • lambda_pos, lambda_vel, lambda_foot = not reported
    Weighting coefficients in Eq. (9) for joint position, velocity, and foot-contact losses; no values are provided, and they shape the generated dance physics.
  • lambda_kpt, lambda_sil = not reported
    Weights in Eq. (14) for keypoint and silhouette alignment in shape optimization; no values given.
  • 80% success threshold in user evaluation = 0.80
    Hand-chosen criterion in Sec. 4.4 for counting a style match in the human study; it is not justified and changes the reported success rates.
assumptions (5)
  • standard math DDPM forward and reverse processes are correctly implemented as in Ho et al.
    Used throughout Sec. 3.1; the paper relies on the standard diffusion formulation without modification.
  • domain assumption AIST++ music and motion pairs provide reliable genre and choreography style labels for contrastive training and classification.
    Sec. 4.1 trains MotionTune and the dance generator on AIST++; if labels are noisy, style control and MSAS are compromised.
  • domain assumption SMPL and 6D rotation pose parameters carry enough information to drive photorealistic video generation.
    Sec. 3.1 and 3.4 use pose as the only bridge between music and video; if pose misses appearance, contact, or expression cues, video quality suffers.
  • domain assumption A music classifier can reliably predict the genre, and the mapped choreography style is acceptable.
    Sec. 3.3's style controller relies on classifier output to set the style; misclassification would produce the wrong dance style.
  • domain assumption The Champ and TikTok video datasets are representative enough to train a video generator that handles any reference individual at any resolution.
    Sec. 4.1 trains the video stage on these datasets and then claims generality beyond them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion." pith.science (2026). https://pith.science/paper/4IP7SKWF

@misc{pith2026250719836,
  author       = {Pith},
  title        = {Pith review of: ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4IP7SKWF}},
  note         = {Machine review of arXiv:2507.19836}
}
read the original abstract

Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance videos that harmonize with both musical rhythm and user-defined choreography styles, limiting their applicability in real-world creative contexts. To address this gap, we introduce ChoreoMuse, a diffusion-based framework that uses SMPL format parameters and their variation version as intermediaries between music and video generation, thereby overcoming the usual constraints imposed by video resolution. Critically, ChoreoMuse supports style-controllable, high-fidelity dance video generation across diverse musical genres and individual dancer characteristics, including the flexibility to handle any reference individual at any resolution. Our method employs a novel music encoder MotionTune to capture motion cues from audio, ensuring that the generated choreography closely follows the beat and expressive qualities of the input music. To quantitatively evaluate how well the generated dances match both musical and choreographic styles, we introduce two new metrics that measure alignment with the intended stylistic cues. Extensive experiments confirm that ChoreoMuse achieves state-of-the-art performance across multiple dimensions, including video quality, beat alignment, dance diversity, and style adherence, demonstrating its potential as a robust solution for a wide range of creative applications. Video results can be found on our project page: https://choreomuse.github.io.

Figures

Figures reproduced from arXiv: 2507.19836 by the authors.

Figure 1
Figure 1. Examples of video frames generated by ChoreoMuse. Given a reference image and a piece of music, ChoreoMuse [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Inference process of ChoreoMuse. Given a piece of [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The training framework of ChoreoMuse. In Stage one, the dance sequence generator is trained using a reference [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Qualitative comparisons between ChoreoMuse and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Generated video frames from ChoreoMuse, demonstrating its ability to process input images of any resolution and [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 52 canonical work pages

  1. [1]

    Emre Aksan, Manuel Kaufmann, and Otmar Hilliges. 2019. Structured prediction helps 3d human motion modelling. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7144–7153

  2. [2]

    Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan. 2023. Person image synthesis via denoising diffusion model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5968–5976

  3. [3]

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. 2023. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22563–22575

  4. [4]

    Judith Butepage, Michael J Black, Danica Kragic, and Hedvig Kjellstrom. 2017. Deep representation learning for human motion prediction and classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6158–6166

  5. [5]

    Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2017. Realtime multi- person 2d pose estimation using part affinity fields. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7291–7299

  6. [6]

    Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022. HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection. In ICASSP 2022-2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing . IEEE, 646–650

  7. [7]

    Yimian Dai, Fabian Gieseke, Stefan Oehmcke, Yiquan Wu, and Kobus Barnard

  8. [8]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186

Show all 75 references
  1. [9]

    Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341 (2020)

  2. [10]

    Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang

  3. [11]

    Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. 2023. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7346–7356

  4. [12]

    Joao P Ferreira, Thiago M Coutinho, Thiago L Gomes, José F Neto, Rafael Azevedo, Renato Martins, and Erickson R Nascimento. 2021. Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio. Computers & Graphics 94 (2021), 11–21

  5. [13]

    Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. 2019. Learning individual styles of conversational gesture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3497–3506

  6. [14]

    Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, and Xinchao Wang. 2023. Tm2d: Bimodality driven 3d dance generation via music-text integration. InProceedings of the IEEE/CVF International Conference on Computer Vision . 9942–9952

  7. [15]

    Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng

  8. [16]

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. 2023. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725 (2023)

  9. [17]

    William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, and Frank Wood. 2022. Flexible diffusion modeling of long videos.Advances in Neural Information Processing Systems 35 (2022), 27953–27965

  10. [18]

    Alejandro Hernandez, Jurgen Gall, and Francesc Moreno-Noguer. 2019. Human motion prediction via spatio-temporal inpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7134–7143

  11. [19]

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303 (2022)

  12. [20]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020), 6840–6851

  13. [21]

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video diffusion models. Advances in Neu- ral Information Processing Systems 35 (2022), 8633–8646

  14. [22]

    Daniel Holden, Jun Saito, and Taku Komura. 2016. A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics 35, 4 (2016), 1–11

  15. [23]

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2022. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868 (2022)

  16. [24]

    Alain Hore and Djemel Ziou. 2010. Image quality metrics: PSNR vs. SSIM. In 2010 20th International Conference on Pattern Recognition . IEEE, 2366–2369

  17. [25]

    Li Hu. 2024. Animate anyone: Consistent and controllable image-to-video syn- thesis for character animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8153–8163

  18. [26]

    Ruozi Huang, Huang Hu, Wei Wu, Kei Sawada, Mi Zhang, and Daxin Jiang

  19. [27]

    Yasamin Jafarian and Hyun Soo Park. 2021. Learning high fidelity depths of dressed humans by watching social media dance videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12753–12762

  20. [28]

    Hsuan-Kai Kao and Li Su. 2020. Temporally guided music-to-body-movement generation. In Proceedings of the 28th ACM International Conference on Multimedia. 147–155

  21. [29]

    Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher- Shlizerman. 2023. Dreampose: Fashion video synthesis with stable diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 22680– 22690

  22. [30]

    Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Hen- schel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. 2023. Text2video-zero: Text-to-image diffusion models are zero-shot video genera- tors. In Proceedings of the IEEE/CVF International Conference on...

  23. [31]

    Jinwoo Kim, Heeseok Oh, Seongjean Kim, Hoseok Tong, and Sanghoon Lee. 2022. A brand new dance partner: Music-conditioned pluralistic dancing controlled by multiple dance genres. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3490–3500

  24. [32]

    Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)

  25. [33]

    Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley. 2020. Panns: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020), 2880–2894

  26. [34]

    Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. 2019. Dancing to music. Advances in Neural Information Processing Systems 32 (2019)

  27. [35]

    Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. 2021. Ai choreog- rapher: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 13401–13412

  28. [36]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015. SMPL: A Skinned Multi-Person Linear Model. ACM Transactions on Graphics 34, 6 (2015)

  29. [37]

    Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion proba- bilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11461–11471

  30. [38]

    Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. 2024. Follow your pose: Pose-guided text-to-video generation using pose-free videos. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 4117–4125

  31. [39]

    Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python.. In SciPy. 18–24

  32. [40]

    Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min. 2023. Conditional image-to-video generation with latent flow diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18444–18455

  33. [41]

    Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32

  34. [42]

    Mathis Petrovich, Michael J Black, and Gül Varol. 2021. Action-conditioned 3D human motion synthesis with transformer VAE. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10985–10995

  35. [43]

    Mathis Petrovich, Michael J Black, and Gül Varol. 2022. Temos: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision. Springer, 480–497

  36. [44]

    Qiaosong Qi, Le Zhuo, Aixi Zhang, Yue Liao, Fei Fang, Si Liu, and Shuicheng Yan

  37. [45]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learni...

  38. [46]

    Xuanchi Ren, Haoran Li, Zijian Huang, and Qifeng Chen. 2020. Self-supervised dance video synthesis conditioned on music. In Proceedings of the 28th ACM International Conference on Multimedia . 46–54

  39. [47]

    Eli Shlizerman, Lucio Dery, Hayden Schoen, and Ira Kemelmacher-Shlizerman

  40. [48]

    Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. 2022. Bailando: 3d dance generation by actor- critic gpt with choreographic memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11050–11059

  41. [49]

    In Proceedings of the 31st ACM International Conference on Multimedia

    Diffdance: Cascaded human motion diffusion model for dance generation. In Proceedings of the 31st ACM International Conference on Multimedia. 1374–1382

  42. [50]

    Guofei Sun, Yongkang Wong, Zhiyong Cheng, Mohan S Kankanhalli, Weidong Geng, and Xiangdong Li. 2020. Deepdance: music-to-dance motion choreography with adversarial learning. IEEE Transactions on Multimedia 23 (2020), 497–509

  43. [51]

    Taoran Tang, Jia Jia, and Hanyang Mao. 2018. Dance with melody: An lstm- autoencoder approach to music-oriented dance synthesis. In Proceedings of the 26th ACM International Conference on Multimedia . 1598–1606

  44. [52]

    Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or

  45. [53]

    Jonathan Tseng, Rodrigo Castellon, and Karen Liu. 2023. Edge: Editable dance generation from music. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 448–458

  46. [54]

    Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. 2018. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717 (2018)

  47. [55]

    Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli

  48. [56]

    Xuanchen Wang, Heng Wang, Dongnan Liu, and Weidong Cai. 2025. Dance Any Beat: Blending Beats with Visuals in Dance Video Generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

  49. [57]

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Transactions on Image Processing 13, 4 (2004), 600–612

  50. [58]

    Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. 2022. Wav2clip: Learning robust audio representations from clip. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 4563–4567

  51. [59]

    Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech an...

  52. [60]

    In European Conference on Computer Vision

    Motionclip: Exposing human motion generation to clip space. In European Conference on Computer Vision . Springer, 358–374

  53. [61]

    Nelson Yalta, Shinji Watanabe, Kazuhiro Nakadai, and Tetsuya Ogata. 2019. Weakly-supervised deep recurrent neural networks for basic dance step genera- tion. In 2019 International Joint Conference on Neural Networks . IEEE, 1–8

  54. [62]

    Zijie Ye, Haozhe Wu, Jia Jia, Yaohua Bu, Wei Chen, Fanbo Meng, and Yanfeng Wang. 2020. Choreonet: Towards music to dance synthesis with choreographic action unit. InProceedings of the 28th ACM International Conference on Multimedia. 744–752

  55. [63]

    Tan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2024. Disco: Disentangled control for realistic human dance generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  56. [64]

    Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5745–5753

  57. [65]

    Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. 2024. Champ: Controllable and consistent human image animation with 3d parametric guidance. In European Conference on Computer Vision. Springer, 145–162

  58. [66]

    Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang, Ming Shao, and Siyu Xia. 2022. Music2dance: Dancenet for music-driven dance generation. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 2 (2022), 1–21

  59. [68]

    Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2024. Magicanimate: Temporally consistent human image animation using diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  60. [71]

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang

  61. [72]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 586–595

  62. [2015]

    In International Conference on Machine Learning

    Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265

  63. [2018]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Audio to body dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7574–7583

  64. [2020]

    arXiv preprint arXiv:2006.06119 (2020)

    Dance revolution: Long-term dance generation with music via curriculum learning. arXiv preprint arXiv:2006.06119 (2020)

  65. [2021]

    InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

    Attentional feature fusion. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 3560–3569

  66. [2022]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Generating diverse and natural 3d human motions from text. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5152– 5161

  67. [2023]

    InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing

    Clap learning audio concepts from natural language supervision. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 1–5

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.