REVIEW 3 major objections 5 minor 55 references
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper introduces HumanTracker, a 153-hour motion-tracking benchmark, and HumanScore, a preference-trained metric that agrees with human judges 90.8% of the time versus 84.1% for the best kinematic diagnostic.
desk verdict Solid benchmark and a plausible preference metric, but the six-annotator single-judgment ground truth and missing learned baselines keep the alignment claim from being fully settled. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is HumanScore, a reward model built from a temporal Transformer. Each frame becomes a 539-dimensional token made of the current reference state plus simulated robot state, actions, measured contact dynamics, root motion, and keypoint kinematics; a padding mask confines attention and mean pooling to real frames, and the pooled token is mapped through an MLP to a scalar reward. Training uses a paired-comparison objective on strict preferences, a symmetric loss for similar pairs, and no cannot-compare pairs; at inference, window rewards pass through a sigmoid and are averaged with frame-length weights to produce a 0-to-100 score. The benchmark side of the paper is a standardized evaluation apparatus: one 29-DoF humanoid, a common simulator entry point, a common success criterion based on pelvis, ankle, and wrist vertical error plus pelvis rotation, and family-level reporting.
What would settle it
Re-run the preference alignment study with a fresh and larger annotation panel, giving each pair multiple independent labels and first measuring inter-annotator agreement; if HumanScore's agreement with those labels does not beat the best kinematic diagnostic, or if the original six annotators fail to agree with the new panel, the claim that HumanScore predicts human preferences collapses.
Extended reading notes
Core claim
HumanScore is a learned trajectory-level score trained to reproduce human comparisons of synchronized tracking rollouts. On the paper's test set, it reaches an Align Rate of 0.9083 with a 95% bootstrap interval from 0.8736 to 0.9383, while the best individual kinematic diagnostic, keypoint position MAE, reaches 0.8405. HumanScore's advantage is largest in contact-rich motions, and it scores Ground and Highly Dynamic rollouts differently than joint error would predict. The paper pairs this metric with a 153-hour benchmark of optical motion capture, retargeted to a 29-DoF humanoid and organized into four motion families, so that tracking quality can be reported per family instead of as one aggregate number. The authors' conclusion is that a preference-aligned metric and a categorized benchmark together reveal contact and stability failures that kinematic metrics miss.
Load-bearing premise
The central claim rests on treating the pairwise judgments of six expert annotators, with a single label per pair and no measured agreement between annotators, as a reliable and stable ground truth for human perception of tracking quality.
Editorial extensions
If this is right
- Tracking results can be reported by motion family, so a tracker that dominates daily upright motions may still collapse on ground-level transitions; aggregate scores alone hide this distinction.
- A common reference representation, rollout accounting, and metric implementation make results across different tracking policies directly comparable rather than confounded by test-set differences.
- HumanScore can be read alongside completion rate and joint error to rank trackers by perceived quality, catching foot sliding and mistimed contacts that pose error misses.
- Longer evaluation windows, up to five seconds, improve alignment with human judgment, indicating that trajectory-level evaluation is needed rather than per-frame error.
- The text-labeled 25K-clip benchmark supports fine-grained failure diagnosis in contact-rich and highly dynamic regimes.
Reading between the lines
- A natural next test is applying HumanScore to real-hardware rollout trajectories; because its current input includes privileged simulator contact and force features, a real-robot study requires an estimated feature set, a boundary the paper itself flags.
- Scaling the preference pipeline from six expert annotators to a larger, more diverse panel would test whether the 90.8% alignment is stable; a plausible outcome is that crowd preferences are noisier but still favor the same contact and stability cues.
- Using HumanScore as a reinforcement-learning reward for tracking policies could improve perceived quality directly, but the paper warns that unregularized optimization of a learned score can be gamed; a controlled experiment with independent human evaluation would settle that.
- The four-family taxonomy invites a failure-regime decomposition: attributing HumanScore to specific events such as impacts, support switches, and recoveries would turn a scalar score into a diagnostic report for controller design.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HumanTracker, a large-scale benchmark for humanoid motion tracking containing approximately 153 hours of optical motion capture from 24 professional performers, organized into four motion families (Daily, Highly Dynamic, Interaction, Ground) with text labels and a standardized MuJoCo evaluation protocol. The paper also proposes HumanScore, a temporal Transformer reward model trained on 6,000 original pairwise preference judgments (mirrored to 12,000 records) collected from six doctoral researchers, and evaluated on a motion-disjoint test split. The central claim is that HumanScore better predicts human preferences than individual kinematic diagnostics such as MPJPE and keypoint-position MAE (Table 4: 90.83% vs 84.05% Align Rate), and that it reveals contact and stability failures that kinematic metrics miss. The paper reports bootstrap confidence intervals clustered by source motion (Table 9), sensitivity analyses of input features and temporal context (Figure 5), and a discussion of limitations including the single-judgment nature of the preference labels.
Significance. If the central claim holds, HumanTracker would be a valuable community resource: it provides a substantially larger and more diverse evaluation suite than the commonly used AMASS test set, a standardized protocol that controls rollout accounting and metric implementation, and a learned perceptual metric that can complement MPJPE-style diagnostics for contact-rich humanoid tracking. The paper has several concrete strengths: the motion-disjoint test split prevents adjacent clips from the same sequence crossing partitions; the bootstrap procedure in Appendix F clusters by source motion; the sensitivity analysis in Figure 5 directly tests the contribution of contact features; and the limitations section (Sec. G) is unusually candid about the scope of the claims, including the single-judgment annotation design and the restriction to one embodiment. These features make the benchmark itself reproducible and the evaluation protocol well specified. However, the headline preference-alignment result currently rests on label reliability and statistical evidence that are not fully established in the manuscript.
major comments (3)
- [Sec. 3.3, Appendix C, Sec. G] The preference ground truth is load-bearing for the central claim, yet it rests on one primary judgment per pair from six doctoral annotators, with no inter-annotator reliability check reported. Both the training labels and the held-out test labels come from the same six-annotator panel, so the Align Rate in Table 4 measures agreement with this specific panel rather than with human preferences broadly. The manuscript itself acknowledges in Sec. G that repeated independent labels were not collected. To support the claim that HumanScore is human-aligned, please report agreement on a double-annotated subset (e.g., Cohen's or Fleiss' kappa), annotator-level agreement rates, and ideally a leave-one-annotator-out training/evaluation analysis. Without such evidence, it is not possible to rule out that the reported 90.83% Align Rate reflects idiosyncratic or systematic biases of the six annotators rather than a general perceptual ground truth.
- [Table 9, Sec. 4.3] The headline numerical advantage of HumanScore over the best kinematic diagnostic is not supported by the reported uncertainty. In Table 9, the 95% bootstrap interval for HumanScore is [87.36, 93.83] and the interval for KPT Position MAE is [79.67, 88.04]; these intervals overlap (between 87.36 and 88.04). Although overlapping marginal intervals do not automatically prove non-significance, the paper does not report a paired test of the difference in Align Rate. To establish that HumanScore is better than the best individual diagnostic, please provide a paired bootstrap or permutation test over source-motion clusters that directly tests the difference, or explicitly state that the current evidence does not show a statistically significant improvement. This is necessary because the 6.8 percentage-point difference is the quantitative basis for the paper's central claim.
- [Abstract, Sec. 4, Fig. 5] The claim that HumanScore 'reveals contact and stability failures that kinematic metrics often miss' is not directly demonstrated in the experiments. Figure 5 shows that removing measured contact features degrades alignment most on Ground, and Table 3 shows that HumanScore ranks SONIC above Humanoid-GPT on Ground despite higher MPJPE. However, there is no qualitative or systematic analysis showing a specific contact or stability failure that HumanScore identifies and MPJPE misses. Please add a case study (e.g., video frames with contact and support annotations) or a quantitative error analysis of disagreements between HumanScore and MPJPE, to substantiate this component of the central claim.
minor comments (5)
- [Table 3] Several entries in Table 3 are missing spaces between numbers, e.g., '97.60.128' and '0.23126.5'; please fix these formatting issues.
- [Sec. 1, Sec. G] The term 'zero-shot generalization' in the Introduction is stronger than what is demonstrated; the current evaluation is motion-level zero-shot only, as Sec. G correctly states. Please qualify the claim in the Introduction to avoid overstating the scope.
- [Sec. 4.1, Table 4] The 'Foot Contact Accuracy' diagnostic in Table 4 is not defined in Sec. 3.2. Please specify how contact agreement is computed (e.g., frame-level binary accuracy, tolerance, and which contact states are compared).
- [Fig. 5] The caption for Figure 5 refers to a 'baseline' that is not explicitly defined in the main text; please state that the baseline is the full model described in Sec. 3.4 and Table 6.
- [Appendix C] Appendix C clarifies that the 12,000 preference records are mirrored variants of 6,000 original annotated pairs. The abstract's '12K motion pairs' could be misread as 12,000 independent judgments; please make the original/mirrored distinction explicit at first mention.
Circularity Check
No significant circularity: HumanScore is trained on preference labels and validated on motion-disjoint, pair-disjoint held-out labels; the central claim is empirical rather than definitional.
full rationale
HumanScore is not derived from the quantities it claims to predict. It is trained with a Bradley–Terry objective (Sec. 3.4) on 6,000 original pairwise labels (mirrored to 12,000 records) collected from six annotators (Sec. 3.3, App. C), and its alignment is measured on an 80/20 motion-disjoint, pair-disjoint test split (Table 4, Table 9). The Table 4 Align Rate is therefore an empirical generalization measurement, not a tautology: the training objective optimizes reward differences on training pairs, while the reported metric is a family-averaged agreement rate on held-out strict comparisons. No fitted parameter is renamed as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work. The only self-citations (e.g., Humanoid-GPT [29], LIMMT [8]) are baselines or related work, not load-bearing justifications for the central claim. The acknowledged limitation in Sec. G — one primary judgment per pair, no repeated independent labels, and the same annotator pool for train and test — affects how broadly the 'human preference' claim generalizes, but it is a data-reliability limitation, not circularity. Similarly, the claim that HumanScore 'reveals contact and stability failures that kinematic metrics often miss' is supported by an ablation (Fig. 5a) and by comparison with kinematic diagnostics, not by construction. The paper explicitly flags in Sec. G that 'the preference pool assigns one primary judgment to each pair... does not quantify uncertainty through repeated independent labels'; this is weighed and found to limit external validity, not to make the derivation circular. Therefore no circular step is exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- HumanScore neural network parameters (weights and biases) =
Trained for 20 epochs on 9,600 preference pairs (80% of 12,000 mirrored records) plus similar pairs, with optimizer…
- Preference model hyperparameters (dimension 256, 4 layers, 8 heads, learning rate 1e-4, dropout 0.1, etc.) =
Listed in Table 6
assumptions (5)
- domain assumption The six doctoral annotators' judgments, with one label per pair, are a reliable ground truth for human preference on tracking quality.
- domain assumption Rollouts from GMT, TWIST2, SONIC, and Humanoid-GPT in MuJoCo cover the distribution of tracking rollouts relevant to benchmarking.
- domain assumption General Motion Retargeting plus manual filtering yields physically valid 29-DoF robot references from optical human mocap.
- domain assumption The 539-dimensional simulator-state feature set is sufficient and correctly represents contact and stability information.
- standard math The Bradley-Terry model and symmetric similar-pair loss correctly model pairwise preferences.
Cite this review
Pith. "Pith review of HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark." pith.science (2026). https://pith.science/paper/X7MOAA23
@misc{pith2026260813555,
author = {Pith},
title = {Pith review of: HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark},
year = {2026},
howpublished = {\url{https://pith.science/paper/X7MOAA23}},
note = {Machine review of arXiv:2608.13555}
}
read the original abstract
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Reference graph
Works this paper leans on
-
[1]
Retargeting matters: General motion retargeting for humanoid motion tracking
Joao Pedro Araujo, Yanjie Ze, Pei Xu, Jiajun Wu, and C Karen Liu. Retargeting matters: General motion retargeting for humanoid motion tracking. ArXiv preprint, abs/2510.02252, 2025.https://ar xiv.org/abs/2510.02252
arXiv 2025
-
[2]
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
Léore Bensabath, Mathis Petrovich, and Gül Varol. A cross-dataset study for text-based 3d human motion retrieval. volume abs/2405.16909, 2024. https://arxiv.org/abs/2405.16909
work page Pith review arXiv 2024
-
[3]
Rank analysis of incomplete block designs: I
Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons.Biometrika, 39(3/4):324–345,
-
[4]
Gmt: General motion tracking for humanoid whole-body control.ArXiv preprint, abs/2506.14770, 2025
Zixuan Chen, Mazeyu Ji, Xuxin Cheng, Xuanbin Peng, Xue Bin Peng, and Xiaolong Wang. Gmt: General motion tracking for humanoid whole-body control.ArXiv preprint, abs/2506.14770, 2025. ht tps://arxiv.org/abs/2506.14770
arXiv 2025
-
[5]
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep re- inforcement learning from human preferences. In Isabelle Guyon, Ulrike von Luxburg, Samy Ben- gio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors,Ad- vances in Neural Information Processing Systems 30: Annual Conference on Ne...
work page 2017
-
[6]
Learning multi-modal whole-body control for real-world humanoid robots
Pranay Dugar, Aayam Shrestha, Fangzhou Yu, Bart van Marum, and Alan Fern. Learning multi-modal whole-body control for real-world humanoid robots. InProceedings of the AAAI Symposium Series, vol- ume 7, pages 650–657, 2025
work page 2025
-
[7]
Humanplus: Humanoid shadowing and imitation from humans
Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wet- zstein, and Chelsea Finn. Humanplus: Humanoid shadowing and imitation from humans. 270:2828– 2844, 2024
work page 2024
-
[8]
LIMMT: Less is more for motion tracking, 2026
Yu Guan, Zekun Qi, Chenghuai Lin, Xuchuan Chen, Dairu Liu, Wenyao Zhang, Jilong Wang, Xinqiang Yu, He Wang, and Li Yi. LIMMT: Less is more for motion tracking, 2026. https://arxiv.org/abs/26 06.06953
work page 2026
Show all 55 references
-
[9]
Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoper- ation and learning.ArXiv preprint, abs/2406.08858, 2024.https://arxiv.org/abs/2406.08858
Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoper- ation and learning.ArXiv preprint, abs/2406.08858, 2024.https://arxiv.org/abs/2406.08858
2024 arXiv
-
[10]
Learning human-to-humanoid real-time whole-body teleopera- tion
Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human-to-humanoid real-time whole-body teleopera- tion. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8944–8951. IEEE, 2024
2024
-
[11]
Switch-justdance: Benchmarking whole body motion tracking policies using a commercial console game.arXiv preprint arXiv:2511.17925, 2025
Jeonghwan Kim, Wontaek Kim, Yidan Lu, Jin Cheng, Fatemeh Zargarbashi, Zicheng Zeng, Zekun Qi, Zhiyang Dou, Nitish Sontakke, Donghoon Baek, et al. Switch-justdance: Benchmarking whole body motion tracking policies using a commercial console game.arXiv preprint arXiv:2511.17925, 2025
2025 arXiv
-
[12]
Phuma: Physically-grounded humanoid lo- comotion dataset.ArXiv preprint, abs/2510.26236, 2025.https://arxiv.org/abs/2510.26236
Kyungmin Lee, Sibeen Kim, Minho Park, Hyunse- ung Kim, Dongyoon Hwang, Hojoon Lee, and Jaegul Choo. Phuma: Physically-grounded humanoid lo- comotion dataset.ArXiv preprint, abs/2510.26236, 2025.https://arxiv.org/abs/2510.26236
2025 arXiv
-
[13]
Robore- ward: General-purpose vision-language reward mod- els for robotics.ArXiv preprint, abs/2601.00675, 2026.https://arxiv.org/abs/2601.00675
Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, and Chelsea Finn. Robore- ward: General-purpose vision-language reward mod- els for robotics.ArXiv preprint, abs/2601.00675, 2026.https://arxiv.org/abs/2601.00675
2026
-
[14]
Object motion guided human motion synthesis.ACM Trans- actions on Graphics (TOG), 42(6):1–11, 2023
Jiaman Li, Jiajun Wu, and C Karen Liu. Object motion guided human motion synthesis.ACM Trans- actions on Graphics (TOG), 42(6):1–11, 2023
2023
-
[15]
Towards motion turing test: Evalu- ating human-likeness in humanoid robots.ArXiv preprint, abs/2603.06181, 2026.https://arxiv.or g/abs/2603.06181
Mingzhe Li, Mengyin Liu, Zekai Wu, Xincheng Lin, Junsheng Zhang, Ming Yan, Zengye Xie, Chang- wang Zhang, Chenglu Wen, Lan Xu, Siqi Shen, and Cheng Wang. Towards motion turing test: Evalu- ating human-likeness in humanoid robots.ArXiv preprint, abs/2603.06181, 2026.https://arx...
2026
-
[16]
Clone: Closed-loop whole-body humanoid teleoperation for long-horizon tasks.ArXiv preprint, abs/2506.08931, 2025.https://arxiv.org/abs/2506.08931
Yixuan Li, Yutang Lin, Jieming Cui, Tengyu Liu, Wei Liang, Yixin Zhu, and Siyuan Huang. Clone: Closed-loop whole-body humanoid teleoperation for long-horizon tasks.ArXiv preprint, abs/2506.08931, 2025.https://arxiv.org/abs/2506.08931
2025 arXiv
-
[17]
Omnitrack: Gen- eral motion tracking via physics-consistent refer- ence.ArXiv preprint, abs/2602.23832, 2026
Yuhan Li, Peiyuan Zhi, Yunshen Wang, Tengyu Liu, Sixu Yan, Wenyu Liu, Xinggang Wang, Baox- iong Jia, and Siyuan Huang. Omnitrack: Gen- eral motion tracking via physics-consistent refer- ence.ArXiv preprint, abs/2602.23832, 2026. https: //arxiv.org/abs/2602.23832
2026
-
[18]
Huang, Luke Zettlemoyer, Dieter Fox, Yu Xiang, Anqi Li, Andreea Bobu, Abhishek Gupta, Stephen Tu, Erdem Biyik, and Jesse Zhang
Anthony Liang, Yigit Korkmaz, Jiahui Zhang, Miny- oung Hwang, Abrar Anwar, Sidhant Kaushik, Aditya Shah, Alex S. Huang, Luke Zettlemoyer, Dieter Fox, Yu Xiang, Anqi Li, Andreea Bobu, Abhishek Gupta, Stephen Tu, Erdem Biyik, and Jesse Zhang. Robometer: Scaling general-purpose r...
2026 arXiv
-
[19]
Beyondmimic: From motion tracking to versatile hu- manoid control via guided diffusion.ArXiv preprint, abs/2508.08241, 2025
Qiayuan Liao, Takara E Truong, Xiaoyu Huang, Guy Tevet, Koushil Sreenath, and C Karen Liu. Beyondmimic: From motion tracking to versatile hu- manoid control via guided diffusion.ArXiv preprint, abs/2508.08241, 2025. https://arxiv.org/abs/25 08.08241
2025 arXiv
-
[20]
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large-scale 3d expressive whole-body human motion dataset. In Alice Oh, Tristan Nau- mann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors,Advances in N...
2023
-
[21]
Smpl: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi-person linear model. InSeminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023
2023
-
[22]
Perpetual hu- manoid control for real-time simulated avatars
Zhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani, and Weipeng Xu. Perpetual hu- manoid control for real-time simulated avatars. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 10861–10870. IEEE, 2023. doi: 10.1...
2023
-
[23]
Kitani, and Weipeng Xu
Zhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang, Kris M. Kitani, and Weipeng Xu. Universal humanoid motion representations for physics-based control. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 20...
2024
-
[24]
Sonic: Supersizing motion tracking for natural humanoid whole-body control.ArXiv preprint, abs/2511.07820, 2025.https://arxiv.org/abs/2511.07820
Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Sirui Chen, Fernando Castañeda, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, et al. Sonic: Supersizing motion tracking for natural humanoid whole-body control.ArXiv preprint, abs/2511.07820, 2025.https://arxiv.org/abs/2511.07820
2025 arXiv
-
[25]
LIV:Language- image representations and rewards for robotic con- trol
Yecheng Jason Ma, Vikash Kumar, Amy Zhang, Os- bertBastani, andDineshJayaraman. LIV:Language- image representations and rewards for robotic con- trol. InProceedings of the 40th International Confer- ence on Machine Learning, volume 202 ofProceedings of Machine Learning Researc...
2023
-
[26]
Troje, Gerard Pons-Moll, and Michael J
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: archive of motion capture as surface shapes. In2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 544...
2019
-
[27]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Pe- ter Welinder, Paul F. Christiano, Jan Le...
2022
-
[28]
Deepmimic: Example-guided deep reinforcement learning of physics-based charac- ter skills.ACM Transactions On Graphics (TOG), 37(4):1–14, 2018
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne. Deepmimic: Example-guided deep reinforcement learning of physics-based charac- ter skills.ACM Transactions On Graphics (TOG), 37(4):1–14, 2018
2018
-
[29]
Humanoid-gpt: Scaling data and struc- ture for zero-shot motion tracking.ArXiv preprint, abs/2606.03985, 2026
Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin, Yunrui Lian, Sikai Liang, Zhikai Zhang, Yu Guan, Ji- long Wang, Wenyao Zhang, Xinqiang Yu, He Wang, and Li Yi. Humanoid-gpt: Scaling data and struc- ture for zero-shot motion tracking.ArXiv preprint, abs/2606.03985, 2026. https...
2026 arXiv
-
[30]
Hu- manoid locomotion as next token prediction
Ilija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran, Sarthak Kamat, Trevor Dar- rell, Koushil Sreenath, and Jitendra Malik. Hu- manoid locomotion as next token prediction. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. T...
2024
-
[31]
Exploring text-to-motion gen- eration with human preference
Jenny Sheng, Matthieu Lin, Andrew Zhao, Kevin Pruvost, Yu-Hui Wen, Yangguang Li, Gao Huang, and Yong-Jin Liu. Exploring text-to-motion gen- eration with human preference. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 1888–1899,
-
[33]
Sontakke, Jesse Zhang, Sébastien M
Sumedh A. Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch, Erdem Biyik, Dorsa Sadigh, Chelsea Finn, and Laurent Itti. Roboclip: One demonstration is enough to learn robot policies. In Advances in Neural Information Processing Systems, volume 36, pages 55681–55693, ...
2023
-
[34]
https://openaccess.thecvf.com/content/ CVPR2024W/HuMoGen/html/Sheng_Exploring_Tex t-to-Motion_Generation_with_Human_Preferen ce_CVPRW_2024_paper.html
-
[35]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, ...
2017
-
[36]
Aligning human motion generation with human perceptions
Haoru Wang, Wentao Zhu, Luyi Miao, Yishu Xu, Feng Gao, Qi Tian, and Yizhou Wang. Aligning human motion generation with human perceptions. InInternational Conference on Learning Represen- tations, 2025. https://proceedings.iclr.cc/pape r_files/paper/2025/hash/c129741a2451e5fefe...
2025
-
[37]
Robo-dopamine: General pro- cess reward modeling for high-precision robotic ma- nipulation.ArXiv preprint, abs/2512.23703, 2025
Huajie Tan, Sixiang Chen, Yijie Xu, Zixiao Wang, Yuheng Ji, Cheng Chi, Yaoxu Lyu, Zhongxia Zhao, Xiansheng Chen, Peterson Co, Shaoxuan Xie, Guo- cai Yao, Pengwei Wang, Zhongyuan Wang, and Shanghang Zhang. Robo-dopamine: General pro- cess reward modeling for high-precision robo...
2025
-
[38]
Iterative preference learning from human feedback: Bridging theory and practice for RLHF under kl- constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang. Iterative preference learning from human feedback: Bridging theory and practice for RLHF under kl- constraint. InForty-first International Conference on Machine Learning, ICML 2024, Vie...
2024
-
[39]
Collision-free humanoid traversal in cluttered indoor scenes, 2026.https: //arxiv.org/abs/2601.16035
Han Xue, Sikai Liang, Zhikai Zhang, Zicheng Zeng, Yun Liu, Yunrui Lian, Jilong Wang, Qingtao Liu, Xuesong Shi, and Li Yi. Collision-free humanoid traversal in cluttered indoor scenes, 2026.https: //arxiv.org/abs/2601.16035
2026
-
[40]
Kungfubot: Physics-based humanoid whole-body control for learning highly- dynamicskills.ArXiv preprint, abs/2506.12851, 2025
Weiji Xie, Jinrui Han, Jiakun Zheng, Huanyu Li, Xinzhe Liu, Jiyuan Shi, Weinan Zhang, Chenjia Bai, and Xuelong Li. Kungfubot: Physics-based humanoid whole-body control for learning highly- dynamicskills.ArXiv preprint, abs/2506.12851, 2025. https://arxiv.org/abs/2506.12851
2025 arXiv
-
[41]
Omnire- target: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction.ArXiv preprint, abs/2509.26633, 2025
Lujie Yang, Xiaoyu Huang, Zhen Wu, Angjoo Kanazawa, Pieter Abbeel, Carmelo Sferrazza, C Karen Liu, Rocky Duan, and Guanya Shi. Omnire- target: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction.ArXiv preprint, abs/2509.26633,...
2025 arXiv
-
[42]
Unitracker: Learning universal whole-body motion tracker for humanoid robots.ArXiv preprint, abs/2507.07356, 2025.https://arxiv.org/abs/2507.07356
Kangning Yin, Weishuai Zeng, Ke Fan, Minyue Dai, Zirui Wang, Qiang Zhang, Zheng Tian, Jingbo Wang, Jiangmiao Pang, and Weinan Zhang. Unitracker: Learning universal whole-body motion tracker for humanoid robots.ArXiv preprint, abs/2507.07356, 2025.https://arxiv.org/abs/2507.07356
2025
-
[43]
I-ctrl: Imitation to control humanoid robots through constrained reinforcement learning
Yashuai Yan, Esteve Valls Mascaro, Tobias Egle, and Dongheui Lee. I-ctrl: Imitation to control humanoid robots through constrained reinforcement learning. ArXiv preprint, abs/2405.08726, 2024.https://ar xiv.org/abs/2405.08726
2024 arXiv
-
[44]
TWIST2: Scal- able, portable, and holistic humanoid data collec- tion system.ArXiv preprint, abs/2511.02832, 2025
Yanjie Ze, Siheng Zhao, Weizhuo Wang, Angjoo Kanazawa, Rocky Duan, Pieter Abbeel, Guanya Shi, Jiajun Wu, and C Karen Liu. TWIST2: Scal- able, portable, and holistic humanoid data collec- tion system.ArXiv preprint, abs/2511.02832, 2025. https://arxiv.org/abs/2511.02832
2025
-
[46]
Twist: Teleoperated whole-body imitation system
Yanjie Ze, Zixuan Chen, JoÃG, o Pedro AraÚjo, Zi- ang Cao, Xue Bin Peng, Jiajun Wu, and C Karen Liu. Twist: Teleoperated whole-body imitation system. ArXiv preprint, abs/2505.02833, 2025.https://ar xiv.org/abs/2505.02833
2025 arXiv
-
[47]
Proxycap: Real-time monocular full-body capture in world space via human-centric proxy-to- motion learning
Yuxiang Zhang, Hongwen Zhang, Liangxiao Hu, Ji- ajun Zhang, Hongwei Yi, Shengping Zhang, and Yebin Liu. Proxycap: Real-time monocular full-body capture in world space via human-centric proxy-to- motion learning. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognitio...
2024
-
[48]
Freemotion: Mocap-free human motion synthesis with multimodal large language models
Zhikai Zhang, Yitang Li, Haofeng Huang, Mingx- ian Lin, and Li Yi. Freemotion: Mocap-free human motion synthesis with multimodal large language models. InEuropean Conference on Computer Vi- sion, pages 403–421. Springer, 2024
2024
-
[49]
Motion-x++: A large-scale multimodal 3d whole-body human motion dataset.ArXiv preprint, abs/2501.05098, 2025
Yuhong Zhang, Jing Lin, Ailing Zeng, Guanlin Wu, ShunlinLu, YurongFu, YuanhaoCai, RuimaoZhang, Haoqian Wang, and Lei Zhang. Motion-x++: A large-scale multimodal 3d whole-body human motion dataset.ArXiv preprint, abs/2501.05098, 2025. ht tps://arxiv.org/abs/2501.05098
2025 arXiv
-
[50]
Learning athletic humanoid tennis skills from imperfect human motion data.arXiv preprint arXiv:2603.12686, 2026
Zhikai Zhang, Haofei Lu, Yunrui Lian, Ziqing Chen, Yun Liu, Chenghuai Lin, Han Xue, Zicheng Zeng, Zekun Qi, Shaolin Zheng, et al. Learning athletic humanoid tennis skills from imperfect human motion data.arXiv preprint arXiv:2603.12686, 2026
2026
-
[51]
Physics-based motion imitation with adversarial differential dis- criminators
Ziyu Zhang, Sergey Bashkirov, Dun Yang, Yi Shi, Michael Taylor, and Xue Bin Peng. Physics-based motion imitation with adversarial differential dis- criminators. InProceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–12. ACM,
2025
-
[52]
Resmimic: From general motion tracking to hu- manoid whole-body loco-manipulation via residual learning.ArXiv preprint, abs/2510.05070, 2025
Siheng Zhao, Yanjie Ze, Yue Wang, C Karen Liu, Pieter Abbeel, Guanya Shi, and Rocky Duan. Resmimic: From general motion tracking to hu- manoid whole-body loco-manipulation via residual learning.ArXiv preprint, abs/2510.05070, 2025. https://arxiv.org/abs/2510.05070. A Implement...
-
[54]
Track any motions under any disturbances.ArXiv preprint, abs/2509.13833, 2025
Zhikai Zhang, Jun Guo, Chao Chen, Jilong Wang, Chenghuai Lin, Yunrui Lian, Han Xue, Zhenrong Wang, Maoqi Liu, Jiangran Lyu, et al. Track any motions under any disturbances.ArXiv preprint, abs/2509.13833, 2025. https://arxiv.org/abs/25 09.13833
2025
-
[191]
https://doi.org/10.1109/CVPR52733.2024 .00191
2024
-
[1952]
doi: 10.1093/biomet/39.3-4.324
-
[2024]
https://openreview.net/forum?id=OrOd8P xOO2
-
[2025]
https: //doi.org/10.1145/3757377.3763819
doi: 10.1145/3757377.3763819. https: //doi.org/10.1145/3757377.3763819
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.