REVIEW 3 major objections 5 minor 1 cited by
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that decoding discrete motion tokens can be reframed as a conditional generation problem, and that a rectified-flow decoder operating in the raw continuous motion space yields smoother, more natural motion than the…
desk verdict Useful decoder replacement with real FID gains on high-capacity tokenizers, but the 'without compromising faithfulness' claim is contradicted by the paper's own tables and needs scoping. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two components carry the method. (1) Condition Projection: each token zt covering q frames is repeated q times, stacked into a vector, linearly projected, and unstacked into q frame-wise conditioning vectors, preserving the temporal correspondence between tokens and frames. (2) Rectified Flow Decoder: a conditional flow-matching model vθ(xt, t, C) trained with the least-squares objective min E||(X1 − X0) − v(Xt, t, C)||², where Xt = tX1 + (1 − t)X0, so that inference is an ODE integration from Gaussian noise to motion, conditioned frame-wise on C. Training on 64-frame sliding windows (rather than full 196-frame sequences) and without attention in the U-Net backbone is what lets the decoder generalize to unseen token sequences in stage 2.
What would settle it
Measure how much token information survives the Condition Projection by training a small linear classifier that predicts the original codebook index from the projected features C on held-out motions. If classification accuracy falls far below the accuracy achievable from the raw codebook embeddings themselves, the projection is the faithfulness bottleneck, and the observed R-Precision drop for T2M-GPT is explained by information loss rather than by anything the flow decoder does.
Extended reading notes
Core claim
The central discovery is a decoder-agnostic upgrade path for discrete motion generation. By repeating each discrete token q times and applying a learned linear projection, DisCoRD obtains frame-wise conditioning features that are concatenated channel-wise with the noisy motion in a conditional flow-matching objective dxt = vθ(xt, t, C) dt, learned from sliding windows of 64 frames. At inference, tokens predicted by any pretrained discrete model are projected to C and integrated with Euler steps from Gaussian noise, yielding continuous-motion output. The paper argues this is why discrete methods can finally match continuous ones on naturalness: 0.032 FID on HumanML3D and 0.169 on KIT-ML (best among compared methods), with sJPE reduced by up to 25% in reconstruction; faithfulness holds when the tokenizer uses residual quantization, while a vanilla VQ-VAE (T2M-GPT) shows a small R-Precision drop.
Load-bearing premise
The learned linear projection that expands each token into q frame-wise features must preserve every detail the decoder needs; if it discards token-specific information, no amount of flow-model capacity can restore faithfulness.
Editorial extensions
If this is right
- Replacing the decoder of MoMask with DisCoRD improves reconstruction FID from 0.019 to 0.011 on HumanML3D and generation FID from 0.045 to 0.032, with sJPE dropping about 25%, while R-Precision stays level or improves slightly.
- On KIT-ML, the same swap takes generation FID from 0.204 to 0.169 for MoMask and from 0.718 to 0.541 for T2M-GPT, showing the gain transfers to the smaller, noisier dataset.
- The method carries over to co-speech gesture (FGD 5.21→4.83 for ProbTalk, 74.88→43.58 for TalkSHOW) and music-to-dance (Distk and Distg move toward the ground-truth spread).
- Decoding speed at 16 Euler steps is on par with MoMask's one-step decoder (0.221s vs 0.244s per batch), and with 2 steps DisCoRD is faster while keeping FID 0.034 and sJPE competitive.
- Residual-quantization levels are the faithfulness lever: as MoMask's RQ level rises from R0 to R5, DisCoRD's R-Precision gain over the baseline grows from −2.8% to +0.6% at Top-1, indicating richer tokens are decoded more faithfully.
Reading between the lines
- If the condition projection is the true bottleneck, a non-linear upsampler (e.g., a small transformer decoder over tokens) is a testable replacement that could remove the small R-Precision losses seen with vanilla VQ-VAEs while keeping the flow decoder's naturalness gains.
- The sJPE metric is not limited to motion: any temporally dense generative model (video, audio, facial animation) where FID-like distributional metrics miss per-sample jitter could adopt the symmetric jerk decomposition as a cheap diagnostic.
- The combination of discrete token prediction for faithfulness and flow decoding for naturalness suggests a new default architecture for conditional motion generation, where the tokenizer is the main remaining quality lever; pushing codebook size or residual depth may yield further monotone gains.
- One concrete test of the paper's causal story: train DisCoRD on tokens whose frame correspondence is artificially scrambled (interleave tokens across time). If FID stays good but sJPE degrades, the method truly relies on the temporal alignment the Condition Projection preserves; if not, the projection's temporal structure matters less than claimed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DisCoRD, a decoder replacement for discrete motion-generation models. Instead of decoding discrete tokens with a feed-forward VQ-VAE decoder, the method repeats each token to frame resolution, projects it to conditioning features, and trains a conditional rectified-flow model to map Gaussian noise to raw motion conditioned on those features. The authors evaluate DisCoRD on motion reconstruction and text-to-motion generation on HumanML3D and KIT-ML, and on co-speech gesture and music-to-dance generation, reporting FID improvements over T2M-GPT, MMM, BAMM, and MoMask while claiming that faithfulness is preserved. They also introduce a new metric, symmetric Jerk Percentage Error (sJPE), designed to detect frame-wise noise and under-reconstruction.
Significance. If the claims are properly scoped, DisCoRD is a useful and non-invasive contribution: it can be attached to any existing discrete motion generator, and the reported FID gains for MoMask and BAMM are substantial and internally consistent. The paper also addresses a real measurement gap by proposing a sample-wise naturalness metric and by providing a user study linking sJPE to human judgment, which is a constructive step beyond relying only on FID and MPJPE. The main weakness is that the abstract and contribution statements overclaim faithfulness preservation, because the paper's own tables show a faithfulness loss for low-capacity tokenizers and a universal increase in reconstruction MPJPE. The underlying mechanism is plausible and the empirical evidence for high-capacity tokenizers is strong, but the claims need to be narrowed and the sJPE validation needs more detail.
major comments (3)
- [Abstract and Contribution 1; Tables 1 and 2] The abstract and Contribution 1 state that DisCoRD 'enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings,' but the paper's own results contradict this as stated. In Table 2, T2M-GPT R-Precision drops on HumanML3D from 0.491/0.680/0.775 to 0.476/0.663/0.760 and on KIT-ML from 0.398/0.606/0.729 to 0.382/0.590/0.715; in Table 1, reconstruction MPJPE increases for every baseline (T2M-GPT: 60.0 to 71.5; MMM: 46.9 to 56.8; MoMask: 29.5 to 33.3). The discussion in Section 4.2 itself acknowledges the T2M-GPT faithfulness decline. The claims should be explicitly scoped to tokenizers with sufficient representational capacity, as Table 5 suggests, or the abstract and contributions must be revised to describe the trade-off rather than asserting no compromise.
- [Abstract; Table 2] The abstract's 'state-of-the-art performance, with FID of 0.032 on HumanML3D and 0.169 on KIT-ML' is not supported on KIT-ML, where Table 2 lists ReMoDiffuse with FID 0.155. Please qualify the SOTA claim (for example, as best among discrete-decoder methods, or as best on naturalness under the sJPE criterion) and reconcile the stated numbers with the full comparison table.
- [Section 4.1 and Supplementary Section C.4] The paper's central naturalness improvement is measured primarily by sJPE, a metric proposed in this same paper and designed specifically to penalize the artifacts that DisCoRD targets. The supplementary user study reports a higher Pearson correlation between sJPE and human naturalness scores than between MPJPE and human scores (0.483 versus 0.181), but it does not report the number of participants, number of rated samples, confidence intervals, or the exact exclusion criterion beyond 'lowest 10% of samples in terms of human score standard deviation.' Because the naturalness claim rests heavily on this metric, a fuller validation protocol and statistical reporting are needed before sJPE can serve as the primary evidence.
minor comments (5)
- [Figure 6 caption] The caption contains a typo: 'fasterer' should be 'faster'.
- [Equation (5)] The notation in Equation (5) is unclear about whether jerk is summed over joints or computed per joint; please define the joint aggregation used to obtain a scalar Jpred,t and Jtrue,t.
- [Table 4] The target values 'Distk →(9.780)' and 'Distg →(7.662)' are unexplained; please define the arrow notation and state explicitly that these are ground-truth reference values.
- [Supplementary Table D] The supplementary table reports FIDk/FIDg degradation for TM2D+DisCoRD (23.98/88.74 versus 19.01/20.09), which appears to conflict with the main-text statement that DisCoRD outperforms the baseline on standard metrics; please clarify which supplementary metrics are considered reliable and why they are reported despite the conflict.
- [Supplementary Table B] Several entries in Table B have unresolved citation placeholders ('Fg-T2M [?]', 'M2DM [?]', 'MotionGPT [?]', 'MotionGPT-2 [?]', 'AttT2M [?]'); these should be replaced with proper references.
Circularity Check
No significant circularity: DisCoRD's rectified-flow decoding is benchmarked against external baselines; the only self-referential element is the authors' sJPE metric, which is not part of the training loss and is externally validated with a user study.
full rationale
The claimed derivation chain is independent. DisCoRD trains a conditional rectified flow decoder using the least-squares objective in Eq. (4), conditioning on frame-wise features extracted from pretrained discrete tokens (Section 3.2). Nothing in the paper defines the decoder's output in terms of its own evaluation metric, and no parameter is fitted to the headline FID or R-Precision numbers. The method is evaluated by replacing the decoders of external baselines (T2M-GPT, MMM, MoMask, BAMM, TalkSHOW, ProbTalk, TM2D) with DisCoRD and comparing against standard metrics (FID, R-Precision, MM-Dist, FGD, user studies). The sJPE metric is introduced by the same authors, which creates a mild self-referential flavor, but it is not a training objective and the paper reports a user study in which sJPE correlates more strongly with human naturalness preference than MPJPE, so the evaluation does not reduce to the method by construction. The claim of 'without compromising faithfulness' is weakened by the paper's own Table 2 (T2M-GPT R-Precision drops) and Table 1 (MPJPE increases for all baselines), but that is a scoping/correctness issue, not circularity. No self-citation chain or imported uniqueness theorem is load-bearing; rectified flow, VQ-VAE, and the baseline tokenizers are all external prior work.
Assumptions & free parameters
free parameters (3)
- Inference sampling steps =
16 (default)
- Training window size =
64 frames
- Condition channel dimension =
256
assumptions (5)
- standard math The conditional rectified flow ODE (Eq. 4) transports Gaussian noise to the motion distribution under the learned vector field v_theta.
- domain assumption A learned linear projection of repeated token embeddings into q frame-wise features retains all task-relevant information in each token.
- ad hoc to paper Jerk is a valid proxy for perceived motion naturalness, so the newly defined sJPE score is a meaningful evaluation target.
- domain assumption Training on sliding windows of 64 frames generalizes to full-length sequences and to token sequences unseen in training.
- domain assumption Pretrained discrete tokens retain sufficient information about fine-grained dynamics for natural motion to be recoverable by a better decoder.
Cite this review
Pith. "Pith review of DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding." pith.science (2026). https://pith.science/paper/FSTEDTI5
@misc{pith2026241119527,
author = {Pith},
title = {Pith review of: DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding},
year = {2026},
howpublished = {\url{https://pith.science/paper/FSTEDTI5}},
note = {Machine review of arXiv:2411.19527}
}
read the original abstract
Human motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion but often struggle to adhere to conditioning signals due to high-dimensional complexity and limited training data. To resolve this 'discord' between discrete and continuous representations we introduce DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding, a novel method that leverages rectified flow to decode discrete motion tokens in the continuous, raw motion space. Our core idea is to frame token decoding as a conditional generation task, ensuring that DisCoRD captures fine-grained dynamics and achieves smoother, more natural motions. Compatible with any discrete-based framework, our method enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings. Extensive evaluations demonstrate that DisCoRD achieves state-of-the-art performance, with FID of 0.032 on HumanML3D and 0.169 on KIT-ML. These results establish DisCoRD as a robust solution for bridging the divide between discrete efficiency and continuous realism. Project website: https://whwjdqls.github.io/discord-motion/
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
A semantically aligned latent space plus multi-token cross-attention conditioning sets a new state of the art in text-to-human-motion generation on HumanML3D.
Reference graph
Works this paper leans on
-
[1]
Building nor- malizing flows with stochastic interpolants
Michael S Albergo and Eric Vanden-Eijnden. Building nor- malizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022. 3
arXiv 2022
-
[2]
Okan Arikan and D. A. Forsyth. Interactive motion generation from examples. ACM Trans. Graph., 21(3):483–490, 2002. 2
work page 2002
-
[3]
A robust and sensitive metric for quantify- ing movement smoothness
Sivakumar Balasubramanian, Alejandro Melendez-Calderon, and Etienne Burdet. A robust and sensitive metric for quantify- ing movement smoothness. IEEE transactions on biomedical engineering, 59(8):2126–2136, 2011. 6
work page 2011
-
[4]
Seamless human motion composition with blended positional encodings
German Barquero, Sergio Escalera, and Cristina Palmero. Seamless human motion composition with blended positional encodings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 6
work page 2024
-
[5]
Vighnesh Birodkar, Gabriel Barcik, James Lyon, Sergey Ioffe, David Minnen, and Joshua V Dillon. Sample what you cant compress. arXiv preprint arXiv:2409.02529, 2024. 2
arXiv 2024
-
[6]
Executing your commands via motion diffusion in latent space
Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu. Executing your commands via motion diffusion in latent space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 18000–18010, 2023. 7, 16
work page 2023
-
[7]
M2d2m: Multi-motion generation from text with discrete diffusion models, 2024
Seunggeun Chi, Hyung gun Chi, Hengbo Ma, Nakul Agar- wal, Faizan Siddiqui, Karthik Ramani, and Kwonjoon Lee. M2d2m: Multi-motion generation from text with discrete diffusion models, 2024. 3, 6, 16
work page 2024
-
[8]
Motionlcm: Real-time controllable motion generation via latent consistency model, 2024
Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, and Yansong Tang. Motionlcm: Real-time controllable motion generation via latent consistency model, 2024. 3
work page 2024
Show all 68 references
-
[9]
A review of 3d human pose estimation algo- rithms for markerless motion capture, 2021
Yann Desmarais, Denis Mottet, Pierre Slangen, and Philippe Montesinos. A review of 3d human pose estimation algo- rithms for markerless motion capture, 2021. 2, 6
2021
-
[10]
Tm2d: Bimodality driven 3d dance generation via music-text integration
Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zi- hang Jiang, Xinxin Zuo, Michael Bi Mi, and Xinchao Wang. Tm2d: Bimodality driven 3d dance generation via music-text integration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9942–9952, 20...
2023
-
[11]
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5152–5161, 2022. 1, 3, 5, 12
2022
-
[12]
Momask: Generative masked modeling of 3d human motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. Momask: Generative masked modeling of 3d human motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1900– 1910, 2024. 1, 2, 3, 6, 7, 12, 14, 16
1900
-
[13]
Parametric motion graphs
Rachel Heck and Michael Gleicher. Parametric motion graphs. In Proceedings of the 2007 symposium on Interactive 3D graphics and games, pages 129–136, 2007. 2
2007
-
[14]
Human motion prediction via spatio-temporal in- painting
Alejandro Hernandez, Jurgen Gall, and Francesc Moreno- Noguer. Human motion prediction via spatio-temporal in- painting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7134–7143, 2019. 3
2019
-
[15]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 3
2020
-
[16]
Fleet, Mohammad Norouzi, and Tim Salimans
Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation, 2021. 5
2021
-
[17]
Bad: Bidirectional auto-regressive diffusion for text-to-motion gen- eration
S Rohollah Hosseyni, Ali Ahmad Rahmani, S Jamal Seyed- mohammadi, Sanaz Seyedin, and Arash Mohammadi. Bad: Bidirectional auto-regressive diffusion for text-to-motion gen- eration. arXiv preprint arXiv:2409.10847, 2024. 3
2024 arXiv
-
[18]
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3. 6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelli- gence, 36(7):1325–1339, 2013. 2, 6
2013
-
[19]
Intermask: 3d human interaction generation via collabora- tive masked modeling, 2025
Muhammad Gohar Javed, Chuan Guo, Li Cheng, and Xingyu Li. Intermask: 3d human interaction generation via collabora- tive masked modeling, 2025. 3
2025
-
[20]
Understanding diffusion ob- jectives as the elbo with simple data augmentation
Diederik Kingma and Ruiqi Gao. Understanding diffusion ob- jectives as the elbo with simple data augmentation. Advances in Neural Information Processing Systems, 36, 2024. 3
2024
-
[21]
Auto-encoding variational bayes
Diederik P Kingma. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 3
2013 arXiv
-
[22]
Motion graphs
Lucas Kovar, Michael Gleicher, and Fr´ed´eric Pighin. Motion graphs. In ACM SIGGRAPH 2008 Classes, New York, NY , USA, 2008. Association for Computing Machinery. 2
2008
-
[23]
A review of com- putable expressive descriptors of human motion
Caroline Larboulette and Sylvie Gibet. A review of com- putable expressive descriptors of human motion. In Proceed- ings of the 2nd International Workshop on Movement and Computing, pages 21–28, 2015. 6
2015
-
[24]
Autoregressive image generation using resid- ual quantization, 2022
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. Autoregressive image generation using resid- ual quantization, 2022. 3, 7
2022
-
[25]
Ross, and Angjoo Kanazawa
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. Ai choreographer: Music conditioned 3d dance generation with aist++, 2021. 1, 3, 5, 12
2021
-
[26]
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. 2, 3
2022 arXiv
-
[27]
Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022
Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, and Bo Zheng. Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022. 3
2022
-
[28]
Haiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng, Mingyang Su, You Zhou, Xuefei Zhe, Naoya Iwamoto, Bo Zheng, and Michael J. Black. Emage: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling, 2024. 1, 3
2024
-
[29]
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022. 2, 3
2022 arXiv
-
[30]
Towards variable and coordinated holis- tic co-speech motion generation
Yifei Liu, Qiong Cao, Yandong Wen, Huaiguang Jiang, and Changxing Ding. Towards variable and coordinated holis- tic co-speech motion generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1566–1576, 2024. 1, 3, 7, 12, 16
2024
-
[31]
Diversemotion: Towards diverse human motion generation via discrete diffusion
Yunhong Lou, Linchao Zhu, Yaxiong Wang, Xiaohan Wang, and Yi Yang. Diversemotion: Towards diverse human motion generation via discrete diffusion. arXiv preprint arXiv:2309.01372, 2023. 3
2023 arXiv
-
[32]
Troje, Ger- ard Pons-Moll, and Michael J
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Ger- ard Pons-Moll, and Michael J. Black. Amass: Archive of motion capture as surface shapes, 2019. 2
2019
-
[33]
Eval- uating the quality of a synthesized motion with the fr ´echet motion distance, 2022
Antoine Maiorca, Youngwoo Yoon, and Thierry Dutoit. Eval- uating the quality of a synthesized motion with the fr ´echet motion distance, 2022. 5, 12
2022
-
[34]
Accuracy measures: theoretical and prac- tical concerns
Spyros Makridakis. Accuracy measures: theoretical and prac- tical concerns. International Journal of Forecasting, 9(4): 527–529, 1993. 6
1993
-
[35]
Synergy and synchrony in couple dances, 2024
V ongani Maluleke, Lea M¨uller, Jathushan Rajasegaran, Geor- gios Pavlakos, Shiry Ginosar, Angjoo Kanazawa, and Jitendra Malik. Synergy and synchrony in couple dances, 2024. 3
2024
-
[36]
MacDorman, and Norri Kageki
Masahiro Mori, Karl F. MacDorman, and Norri Kageki. The uncanny valley [from the field]. IEEE Robotics & Automation Magazine, 19(2):98–100, 2012. 2
2012
-
[37]
Generative proxemics: A prior for 3d social interaction from images, 2023
Lea M ¨uller, Vickie Ye, Georgios Pavlakos, Michael Black, and Angjoo Kanazawa. Generative proxemics: A prior for 3d social interaction from images, 2023. 3
2023
-
[38]
Black, and G¨ul Varol
Mathis Petrovich, Michael J. Black, and G¨ul Varol. Action- conditioned 3D human motion synthesis with transformer V AE. In International Conference on Computer Vision (ICCV), 2021. 2
2021
-
[39]
Temos: Generating diverse human motions from textual descriptions
Mathis Petrovich, Michael J Black, and G ¨ul Varol. Temos: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision, pages 480–497. Springer, 2022. 3
2022
-
[40]
Bamm: Bidirectional autoregressive motion model
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee, Srijan Das, and Chen Chen. Bamm: Bidirectional autoregressive motion model. arXiv preprint arXiv:2403.19435, 2024. 2, 3, 7, 16
2024 arXiv
-
[41]
Mmm: Generative masked motion model
Ekkasit Pinyoanuntapong, Pu Wang, Minwoo Lee, and Chen Chen. Mmm: Generative masked motion model. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1546–1555, 2024. 2, 6, 7, 14, 16
2024
-
[42]
The kit motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset. Big data, 4(4):236–252,
-
[43]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[44]
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 5
2022
-
[45]
Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H. Bermano. Human motion diffusion as a generative prior, 2023. 3
2023
-
[46]
Bailando: 3d dance generation by actor-critic gpt with choreographic mem- ory
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. Bailando: 3d dance generation by actor-critic gpt with choreographic mem- ory. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 110...
-
[47]
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 3
2011 arXiv
-
[48]
Kankanhalli, Weidong Geng, and Xiangdong Li
Guofei Sun, Yongkang Wong, Zhiyong Cheng, Mohan S. Kankanhalli, Weidong Geng, and Xiangdong Li. Deepdance: Music-to-dance motion choreography with adversarial learn- ing. IEEE Transactions on Multimedia, 23:497–509, 2021. 3
2021
-
[49]
Grounding multi- modal large language models in actions, 2024
Andrew Szot, Bogdan Mazoure, Harsh Agrawal, Devon Hjelm, Zsolt Kira, and Alexander Toshev. Grounding multi- modal large language models in actions, 2024. 2
2024
-
[50]
Motionclip: Exposing human motion gen- eration to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion gen- eration to clip space. In European Conference on Computer Vision, pages 358–374. Springer, 2022. 3
2022
-
[51]
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2023. 1, 2, 5
2023
-
[52]
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2023. 3, 7, 16
2023
-
[53]
Edge: Editable dance generation from music
Jonathan Tseng, Rodrigo Castellon, and Karen Liu. Edge: Editable dance generation from music. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 448–458, 2023. 3, 5, 12, 15
2023
-
[54]
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information pro- cessing systems, 30, 2017. 3
2017
-
[55]
Neural discrete representation learning, 2018
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning, 2018. 2
2018
-
[56]
What is the best automated metric for text to motion generation? In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023
Jordan V oas, Yili Wang, Qixing Huang, and Raymond Mooney. What is the best automated metric for text to motion generation? In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023. 1
2023
-
[57]
Aligning motion generation with human perceptions
Haoru Wang, Wentao Zhu, Luyi Miao, Yishu Xu, Feng Gao, Qi Tian, and Yizhou Wang. Aligning motion generation with human perceptions. 2024. 6
2024
-
[58]
Ecological validity and the evaluation of avatar facial anima- tion noise
Marta Wilczkowiak, Ken Jakubzak, James Clemoes, Cornelia Treptow, Michaela Porubanova, Kerry Read, Daniel McDuff, Marina Kuznetsova, Sean Rintel, and Mar Gonzalez-Franco. Ecological validity and the evaluation of avatar facial anima- tion noise. In 2024 IEEE Conference on Virt...
2024
-
[59]
Motionllm: Multimodal motion-language learning with large language models
Qi Wu, Yubo Zhao, Yifan Wang, Yu-Wing Tai, and Chi-Keung Tang. Motionllm: Multimodal motion-language learning with large language models. arXiv preprint arXiv:2405.17013 ,
-
[60]
Executing your com- mands via motion diffusion in latent space
Chen Xin, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, Jingyi Yu, and Gang Yu. Executing your com- mands via motion diffusion in latent space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2, 3, 5, 6
2023
-
[61]
Codetalker: Speech-driven 3d facial animation with discrete motion prior
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. Codetalker: Speech-driven 3d facial animation with discrete motion prior. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12780–12790, 2023. 2
2023
-
[62]
Gen- erating holistic 3d human motion from speech
Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J Black. Gen- erating holistic 3d human motion from speech. In CVPR,
-
[63]
Physdiff: Physics-guided human motion diffusion model
Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. Physdiff: Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF international confer- ence on computer vision, pages 16010–16021, 2023. 3
2023
-
[64]
T2m-gpt: Generating human motion from textual de- scriptions with discrete representations, 2023
Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2m-gpt: Generating human motion from textual de- scriptions with discrete representations, 2023. 2, 3, 6, 7, 14, 16
2023
-
[65]
Motiondiffuse: Text- driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondiffuse: Text- driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001, 2022. 2, 3, 7, 16
2022 arXiv
-
[66]
Re- modiffuse: Retrieval-augmented motion diffusion model
Mingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai, Fangzhou Hong, Huirong Li, Lei Yang, and Ziwei Liu. Re- modiffuse: Retrieval-augmented motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 364–373, 2023. 3, 7, 16
2023
-
[67]
person is acting like human monkey
Long Zhao, Sanghyun Woo, Ziyu Wan, Yandong Li, Han Zhang, Boqing Gong, Hartwig Adam, Xuhui Jia, and Ting Liu. ϵ-V AE: Denoising as Visual Decoding.arXiv preprint arXiv:2410.04081, 2024. 2, 5 DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding Supplementar...
2024 arXiv
-
[2023]
1, 3, 5, 7, 12, 15, 16
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.