REVIEW 4 major objections 6 minor 119 references
Towards Embodiment Scaling Laws in Robot Locomotion
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Increasing the number of robot bodies a locomotion policy is trained on improves its ability to control unseen bodies, and this embodiment scaling helps more than adding data on a fixed set of bodies.
desk verdict First ~1k-robot embodiment scaling study with a real in-distribution trend; the law-level claim needs stronger external validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an embodiment-conditioned attention policy built on URMA, a joint-level architecture that handles arbitrary robot morphologies by splitting observations into fixed general features and variable-length per-joint features. A multi-head attention encoder fuses joint observations, weighted by learned joint-description vectors derived from the embodiment descriptor $\phi(e)$, so the same network can output actions for bodies with different joint counts and kinematic properties. The scaling study is carried out on GENBOT-1K, a dataset of 1,012 procedurally generated blueprints varying topology (number of knee joints), geometry (link lengths and sizes), and kinematics (joint limits), with a fixed 20% held-out test set. Training follows a two-stage pipeline: per-embodiment RL experts provide demonstrations, and a single student policy is distilled from them by behavior cloning.
What would settle it
Retrain the scaling curves with test bodies drawn from a separate generator that includes factors the paper holds fixed, such as mass distribution, joint damping, and actuation type, and with joint limits outside the training ranges; if held-out reward no longer rises with training embodiment count, the observed law is an artifact of sampling density inside the generator.
Extended reading notes
Core claim
The paper's central claim is that, for flat-ground proprioceptive locomotion, generalization to held-out robot bodies improves as the number of training embodiments grows, and that this embodiment scaling is not reducible to data scaling. On a fixed test set of 204 procedurally generated robots, held-out reward roughly doubles when the training embodiment fraction rises from 5% to 100% of the generated pool. A control trained on only 5% of the bodies but with four times as many demonstrations per body nearly saturates, which the paper reads as evidence that body diversity, not trajectory count, drives the gain. Training across humanoids, quadrupeds, and hexapods together yields one policy that beats class-only policies on the mixed test set and transfers zero-shot to the Unitree Go2 and H1, including stable adaptation when a real knee joint's range is artificially reduced to 20% of nominal.
Load-bearing premise
The held-out test embodiments come from the same procedural generator as the training bodies, and the real robots are similar in kinematic structure to bodies in the training set, so the measured scaling is interpolation inside one hand-designed distribution rather than generalization across the full space of possible robot bodies.
Editorial extensions
If this is right
- A single locomotion policy trained on a sufficiently diverse set of simulation bodies can be deployed on new robot hardware without per-robot fine-tuning, as demonstrated by the zero-shot real-world transfers.
- When the goal is cross-embodiment generalization, spending a fixed data budget on more robot bodies is more effective than collecting more demonstrations from a few bodies.
- Scaling curves can serve as a planning tool: harder morphology classes, such as humanoids, may require more training embodiments to reach the same level of held-out performance as easier classes.
- The policy's latent representations organize by morphology class and joint count, suggesting that embodiment-aware controllers can adapt to a changed body, such as a restricted joint, by reading its descriptor and applying the nearest learned behavior.
Reading between the lines
- If the scaling trend reflects a general law rather than generator interpolation, the same axis should appear in manipulation and whole-body control, though the curve may shift because those tasks add perceptual variation and contact-rich dynamics beyond morphology.
- A stricter test of the law would adversarially select training bodies to maximize coverage of kinematic extremes; if the curve flattens, the active variable is distribution coverage rather than embodiment count itself.
- The joint-limit adaptation shown on one real robot suggests that a single policy could control modular or reconfigurable robots whose geometry changes between deployments, provided the changes stay within the parameter ranges the generator was built from.
- The data-scaling saturation result implies a practical stopping rule for data collection: once additional trajectories from a fixed robot stop improving held-out performance, the remaining budget is better spent on new morphologies, but this rule is an extrapolation beyond the paper's measured regime.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether increasing the number of training robot embodiments improves generalization to unseen embodiments, a hypothesis the authors term an "embodiment scaling law." Using locomotion as a testbed, they procedurally generate 1,012 robots spanning humanoids, quadrupeds, and hexapods, with variations in topology, geometry, and joint kinematics. They train single-embodiment expert policies with RL and distill them into a single URMA-based policy via behavior cloning, varying the number of training embodiments from 5% to 100% of an 80% training pool and evaluating on a fixed 20% held-out set. They also compare against a data-scaling baseline (C8) that increases trajectories on a fixed 5% embodiment subset, and they demonstrate zero-shot transfer of the full policy to the Unitree Go2 and H1 robots in the real world, including with artificially restricted knee joint limits. The paper reports positive scaling trends in all three morphology classes and in the combined cross-class setting, and concludes that embodiment scaling enables substantially broader generalization than data scaling.
Significance. If the central claim holds, this would be an important first large-scale empirical step toward understanding how embodiment diversity drives generalization in robot learning, with implications for generalist robot policies, adaptive control, and morphology co-design. The study's strengths are its unprecedented scale (1,012 embodiments, 2 trillion simulation steps), the two-stage RL-to-distillation pipeline that makes such scale tractable, the fixed held-out test set design, the honest limitations section, and the genuine zero-shot transfer to two real robots including a constrained-joint deployment. The latent-space analyses strengthen the plausibility of the mechanism. However, the evidence is currently insufficient to establish a "law": the scaling curves are single runs without uncertainty quantification, the test distribution is the same finite procedural generator as the training distribution, and the data-scaling comparison is not matched on total samples or compute. These are fixable with additional experiments and analyses, but they are load-bearing for the headline claims.
major comments (4)
- [Sec. 4.1, Figure 4] The scaling curves C1-C8 are each single training runs with no error bars, confidence intervals, or multiple seeds. The central claims that Jtest increases monotonically with the number of training embodiments, that quadruped and hexapod performance saturates around 100 embodiments, and that humanoid performance "continues to improve steadily" cannot be distinguished from run-to-run variance at this level of evidence. I request at least 3-5 seeds per condition, or a bootstrap/confidence-interval analysis over the test embodiments, to quantify the trend and its saturation behavior.
- [Sec. 4.1, Appendix B.2, Table 6] The held-out test embodiments are sampled from the same discrete procedural generator used to create the training set, whose parameter grid is coarse: thigh and calf length scales take five values, foot size two values, knee-limit scales three values, and topology is varied by knee count in {0,1,2,3}. As the training subset grows from 5% to 100%, a test embodiment is increasingly likely to share all parameter values with some training body, so the observed rise in Jtest may measure nearest-neighbor coverage of a finite grid rather than a generalizable scaling property of embodiment diversity. The out-of-distribution experiments in Appendix D evaluate only the policy trained on the full set, not whether the scaling trend itself survives outside the training grid. To support the claimed law, the scaling curves should be recomputed on a genuinely external family of embodiments, for example from a different generator, from human-designed robots, or using parameter values not present in the training grid.
- [Sec. 4.1, curve C8] The data-scaling comparison is confounded: C8 fixes the embodiment set at 5% and varies the number of trajectories per embodiment, while the embodiment-scaling curves vary the number of embodiments with roughly fixed per-embodiment data. Consequently, the total number of demonstration samples changes along both axes, so the conclusion that "embodiment scaling is essential" is not supported by a matched comparison. Please include a control in which total sample count (or compute) is held constant while the ratio of the number of embodiments to per-embodiment data is varied.
- [Sec. 4.1, Figure 4 caption] The cross-class comparisons (C4 vs C5-C7) are made on unnormalized rewards whose scales differ across morphology classes, as the caption itself acknowledges. The claimed 2-5x improvement in average reward on the combined test set could be dominated by the class(es) with larger-magnitude rewards where single-class policies fail. Please report per-class normalized rewards or a per-class performance table before drawing the conclusion that training across morphologies enables "substantially broader generalization."
minor comments (6)
- [Eq. (3)] The softmax denominator in Eq. (3) is written as a sum over the latent dimension Ld, but the attention normalization should be over the joints J within an embodiment; please correct the equation and clarify the output dimensionality of f_phi(d_j).
- [Sec. 4.2 vs Appendix C.3] Section 4.2 states that the full policy was trained on 817 simulated embodiments, but Table 9 in Appendix C.3 reports a training set of 808 embodiments (278 humanoid + 265 quadruped + 265 hexapod); please reconcile this inconsistency.
- [Appendix E.2] The paragraph on the software-level joint-limit implementation contains a redundant and slightly contradictory pair of sentences: "we introduce a software-level joint-limit layer into the control loop" and then "Instead, we implemented a software-based solution..."; please edit for clarity.
- [Figure 4] The x-axis of Figure 4 mixes two different quantities (proportion of training embodiments for C1-C7 and data scale for C8) on the same axis; please relabel or split the panels to avoid conflating these axes.
- [Abstract and Section 5] The paper uses the term "embodiment scaling laws" in the title, abstract, and conclusion, but no functional form of a law is actually fitted or verified; please consider tempering the terminology to "embodiment scaling trends" unless a quantitative law is established.
- [References] References [40] and [89] both refer to the same paper (GET-Zero); please cite it only once.
Circularity Check
No significant circularity: the central embodiment-scaling claim is measured on a fixed held-out test set and is not reduced to any fitted input or self-citation.
full rationale
The paper's central claim is an empirical scaling measurement: policies distilled from expert demonstrations on randomized subsets of the training embodiments are evaluated on a fixed 20% held-out test set (Eqs. 1-2, Sec. 3, Sec. 4.1). No test-set quantity enters the distillation loss or the RL expert training, and no parameter is fitted to the held-out reward before it is reported as Jtest. The comparison with pure data scaling (C8) is likewise an empirical control, not a construction that forces the conclusion. The use of URMA [41] is an architectural implementation choice whose cited prior work is external and machine-checked only in the loose sense of being a published architecture; it is not load-bearing for the scaling-law claim, since the same URMA student is used across all training-subset sizes and the trend is measured across those sizes. The paper's self-citations (e.g., [41] for the architecture, [42,43] for the two-stage learning paradigm) do not supply the scaling result itself, and no uniqueness theorem or ansatz is imported to forbid alternative explanations. The limitations section candidly states that the procedural generation does not exhaustively cover the design space and that real-world validation is limited to two platforms; this is an external-validity concern about how broadly the trend generalizes, not evidence that the derivation is circular. In short, the observed Jtest curves are genuine measurements on held-out embodiments, so the central result is self-contained with respect to its inputs.
Assumptions & free parameters
free parameters (1)
- Humanoid-specific reward coefficients =
T1=3.0, T2=1.5, T6=43.2, T17=6e-3
assumptions (3)
- domain assumption The procedural generation distribution over topology, geometry, and kinematics is representative of the embodiment variation that matters for locomotion generalization.
- domain assumption The simulator (Isaac Lab, configured as in Appendix A) is a sufficiently faithful model for sim-to-real transfer of the distilled policy.
- domain assumption Behavior cloning from expert RL policies is a valid proxy for optimizing the cross-embodiment objective in Eq. 1.
Cite this review
Pith. "Pith review of Towards Embodiment Scaling Laws in Robot Locomotion." pith.science (2026). https://pith.science/paper/KXKJWMCC
@misc{pith2026250505753,
author = {Pith},
title = {Pith review of: Towards Embodiment Scaling Laws in Robot Locomotion},
year = {2026},
howpublished = {\url{https://pith.science/paper/KXKJWMCC}},
note = {Machine review of arXiv:2505.05753}
}
read the original abstract
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodiment scaling laws, the hypothesis that increasing the number of training embodiments improves generalization to unseen ones, using robot locomotion as a test bed. We procedurally generate ~1,000 embodiments with topological, geometric, and joint-level kinematic variations, and train policies on random subsets. We observe positive scaling trends supporting the hypothesis, and find that embodiment scaling enables substantially broader generalization than data scaling on fixed embodiments. Our best policy, trained on the full dataset, transfers zero-shot to novel embodiments in simulation and the real world, including the Unitree Go2 and H1. These results represent a step toward general embodied intelligence, with relevance to adaptive control for configurable robots, morphology co-design, and beyond.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollár, and R. Girshick. Segment anything. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3992–4003, 2023. doi:10.1109/ ICCV51070.2023.00371
arXiv 2023
-
[2]
B. Wen, W. Yang, J. Kautz, and S. Birchfield. Foundationpose: Unified 6d pose estima- tion and tracking of novel objects. In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024 , pages 17868–17879. IEEE, 2024. doi:10.1109/CVPR52733.2024.01692. URL https://doi.org/10.1109/ CVPR52733.2024.01692
arXiv 2024
-
[3]
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. In 2021 IEEE/CVF Interna- tional Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10- 17, 2021 , pages 9630–9640. IEEE, 2021. doi:10.1109/ICCV48922.2021.00951. URL https://doi.org/10.1109/ICC...
arXiv 2021
-
[4]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P. Huang, S. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jégou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski. Dinov2: Learning robust visual features without supervi-...
2024
-
[5]
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer. Scaling vision transformers. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 , pages 1204–1213. IEEE, 2022. doi:10.1109/CVPR52688.2022. 01179. URL https://doi.org/10.1109/CVPR52688.2022.01179
arXiv 2022
-
[6]
C. Sun, A. Shrivastava, S. Singh, and A. Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 , pages 843–852. IEEE Computer Society, 2017. doi:10.1109/ICCV .2017.97. URLhttps://doi.org/10.1109/ICCV.2017.97
doi:10.1109/iccv 2017
-
[7]
T. Tian, H. Li, B. Ai, X. Yuan, Z. Huang, and H. Su. Diffusion dynamics models with generative state estimation for cloth manipulation. CoRR, abs/2503.11999, 2025. doi:10. 48550/ARXIV .2503.11999. URLhttps://doi.org/10.48550/arXiv.2503.11999
-
[8]
D. Mahajan, R. B. Girshick, V . Ramanathan, K. He, M. Paluri, Y . Li, A. Bharambe, and L. van der Maaten. Exploring the limits of weakly supervised pretraining. In V . Ferrari, M. Hebert, C. Sminchisescu, and Y . Weiss, editors,Computer Vision - ECCV 2018 - 15th Eu- ropean Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part II, volume 112...
doi:10.1007/978-3-03 2018
Show all 119 references
-
[9]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe. Training lan- guage models to follow instruct...
2022
-
[10]
Achiam, S
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[11]
DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y . Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. L...
2025 arXiv
-
[13]
Chowdhery, S
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y . Tay, N. Shazeer, V . Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. ...
2022 arXiv
-
[14]
T. Gao, A. Fisch, and D. Chen. Making pre-trained language models better few-shot learners. In C. Zong, F. Xia, W. Li, and R. Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference o...
-
[15]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Milli- can, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Si...
-
[16]
B. Ai, Y . Wang, Y . Tan, and S. Tan. Whodunit? learning to contrast for authorship attribution. In Y . He, H. Ji, Y . Liu, S. Li, C. Chang, S. Poria, C. Lin, W. L. Buntine, M. Liakata, H. Yan, Z. Yan, S. Ruder, X. Wan, M. Arana-Catania, Z. Wei, H. Huang, J. Wu, M. Day, P. Liu...
2022
-
[17]
Z. Wu, B. Ai, and D. Hsu. Integrating common sense and planning with large language models for room tidying. In RSS 2023 Workshop on Learning for Task and Motion Planning,
2023
-
[18]
Q. Gao, X. Pi, K. Liu, J. Chen, R. Yang, X. Huang, X. Fang, L. Sun, G. Kishore, B. Ai, S. Tao, M. Liu, J. Yang, C.-J. Lai, C. Jin, J. Xiang, B. Huang, D. Danks, H. Su, T. Shu, Z. Ma, L. Qin, and Z. Hu. Do vision-language models have internal world models? towards an atomic eva...
2025
-
[19]
K. Fang, P. Yin, A. Nair, H. Walke, G. Yan, and S. Levine. Generalization with lossy affor- dances: Leveraging broad offline data for learning visuomotor tasks. In K. Liu, D. Kulic, and 11 J. Ichnowski, editors, Conference on Robot Learning, CoRL 2022, 14-18 December 2022, Auc...
2022
-
[20]
Kumar, A
A. Kumar, A. Singh, F. D. Ebert, M. Nakamoto, Y . Yang, C. Finn, and S. Levine. Pre- training for robots: Offline RL enables learning new tasks in a handful of trials. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors, Robotics: Science and Systems XIX, Daegu, Repu...
2023 doi
-
[22]
Ghosh, H
D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, Q. Vuong, T. Xiao, P. R. Sanketi, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy. In D. Kulic, G. Venture, K. E. Bekr...
2024 doi
-
[23]
Brohan, N
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K. Lee, S. Levine, Y . Lu, U. Malla, D. Manj...
2023
-
[24]
Zitkovich, T
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V . Vanhoucke, H. T. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y ...
2023
-
[25]
Intelligence, K
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A....
2025 arXiv
-
[26]
H. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu. RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot. In IEEE Interna- tional Conference on Robotics and Automation, ICRA 2024, Yokohama, Japan, May 13- 17, 2024 , pages 653–660. ...
2024
-
[27]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model. In P. Agrawal, O....
2024
-
[28]
Black, N
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, et al. π0: A vision-language-action flow model for general robot control,
-
[29]
Ebert, Y
F. Ebert, Y . Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. In K. Hauser, D. A. Shell, and S. Huang, editors, Robotics: Science and Systems XVIII, New Yo...
2022 doi
-
[30]
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V . Myers, M. J. Kim, M. Du, A. Lee, K. Fang, C. Finn, and S. Levine. Bridgedata V2: A dataset for robot learning at scale. In J. Tan, M. Toussaint, and K. Darvish, editors,Conference on Robot ...
2023
- [31]
-
[32]
H. Fang, C. Wang, M. Gou, and C. Lu. Graspnet-1billion: A large-scale benchmark for gen- eral object grasping. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020 , pages 11441–11450. Computer Vision Foundatio...
2020
-
[33]
W. Gao, B. Ai, J. Loo, Vinay, and D. Hsu. Intentionnet: Map-lite visual navigation at the kilometre scale, 2024. URL https://arxiv.org/abs/2407.03122
2024 arXiv
-
[34]
B. Ai, Z. Wu, and D. Hsu. Invariance is key to generalization: Examining the role of repre- sentation in sim-to-real transfer for visual navigation. In M. H. Ang Jr and O. Khatib, editors, Experimental Robotics, pages 69–80, Cham, 2024. Springer Nature Switzerland. ISBN 978- 3...
2024
-
[35]
B. Ai, W. Gao, Vinay, and D. Hsu. Deep visual navigation under partial observability. In2022 International Conference on Robotics and Automation, ICRA 2022, Philadelphia, PA, USA, May 23-27, 2022 , pages 9439–9446. IEEE, 2022. doi:10.1109/ICRA46639.2022.9811598. URL https://do...
2022
-
[36]
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y . Liu, I. Misra, S. Chintala, and L. Pinto. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023
2023 arXiv
-
[37]
Mandlekar, S
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y . S. Narang, L. Fan, Y . Zhu, and D. Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations. In J. Tan, M. Toussaint, and K. Darvish, editors,Conference on Robot Learning, CoRL 2023, 6-9 ...
2023
-
[38]
Dasari, F
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn. Robonet: Large-scale multi-robot learning. In L. P. Kaelbling, D. Kragic, and K. Sug- iura, editors, 3rd Annual Conference on Robot Learning, CoRL 2019, Osaka, Japan, Octo- ber...
2019
-
[39]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...
2024
- [40]
-
[41]
Bohlinger, G
N. Bohlinger, G. Czechmanowski, M. Krupka, P. Kicki, K. Walas, J. Peters, and D. Tateo. One policy to run them all: an end-to-end learning approach to multi-embodiment locomotion. Conference on Robot Learning, 2024
2024
-
[42]
Z. Jia, X. Li, Z. Ling, S. Liu, Y . Wu, and H. Su. Improving policy optimization with generalist-specialist learning. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvári, G. Niu, and S. Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, ...
2022
-
[43]
W. Wan, H. Geng, Y . Liu, Z. Shan, Y . Yang, L. Yi, and H. Wang. Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist- specialist learning. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, ...
2023
-
[44]
R. Zhu, T. Dai, and O. Celiktutan. Cross domain policy transfer with effect cycle-consistency. In 2024 IEEE International Conference on Robotics and Automation. IEEE Explore, 2024
2024
-
[45]
Y . Chen, Y . Chen, Z. Hu, T. Yang, C. Fan, Y . Yu, and J. Hao. Learning action-transferable policy with action embedding. arXiv preprint arXiv:1909.02291, 2019
1909 arXiv
-
[46]
Hu and G
Y . Hu and G. Montana. Skill transfer in deep reinforcement learning under morphological heterogeneity. arXiv preprint arXiv:1908.05265, 2019
1908 arXiv
-
[47]
X. Liu, D. Pathak, and D. Zhao. Meta-evolve: Continuous robot evolution for one-to-many policy transfer. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[48]
T. Wang, R. Liao, J. Ba, and S. Fidler. Nervenet: Learning structured policy with graph neural networks. In International Conference on Learning Representations, 2018
2018
-
[49]
Huang, I
W. Huang, I. Mordatch, and D. Pathak. One policy to control them all: Shared mod- ular policies for agent-agnostic control. In Proceedings of the 37th International Con- ference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event , volume 119 of Proceedings of Machi...
2020
-
[50]
Trabucco, M
B. Trabucco, M. Phielipp, and G. Berseth. Anymorph: Learning transferable polices by inferring agent morphology. InInternational Conference on Machine Learning, pages 21677– 21691. PMLR, 2022
2022
-
[51]
Furuta, Y
H. Furuta, Y . Iwasawa, Y . Matsuo, and S. S. Gu. A system for morphology-task generalization via unified representation and behavior distillation. InThe Eleventh International Conference on Learning Representations, 2022
2022
-
[52]
D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine. Gnm: A general navigation model to drive any robot. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 7226–7233. IEEE, 2023
2023
-
[53]
W. Song, H. Zhao, P. Ding, C. Cui, S. Lyu, Y . Fan, and D. Wang. Germ: A generalist robotic model with mixture-of-experts for quadruped robot. arXiv preprint arXiv:2403.13358, 2024. 15
2024 arXiv
-
[54]
Doshi, H
R. Doshi, H. R. Walke, O. Mees, S. Dasari, and S. Levine. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. In 8th Annual Conference on Robot Learning, 2024
2024
-
[55]
Shafiee, G
M. Shafiee, G. Bellegarda, and A. Ijspeert. Manyquadrupeds: Learning a single locomotion policy for diverse quadruped robots. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 3471–3477. IEEE, 2024
2024
-
[56]
Eftekhar, L
A. Eftekhar, L. Weihs, R. Hendrix, E. Caglar, J. Salvador, A. Herrasti, W. Han, E. VanderBil, A. Kembhavi, A. Farhadi, et al. The one ring: a robotic indoor navigation generalist. arXiv preprint arXiv:2412.14401, 2024
2024 arXiv
-
[57]
G. Feng, H. Zhang, Z. Li, X. B. Peng, B. Basireddy, L. Yue, Z. Song, L. Yang, Y . Liu, K. Sreenath, and S. Levine. Genloco: Generalized locomotion controllers for quadrupedal robots. In K. Liu, D. Kulic, and J. Ichnowski, editors, Conference on Robot Learning, CoRL 2022, 14-18...
2022
-
[58]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimiza- tion algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[59]
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning robust perceptive locomotion for quadrupedal robots in the wild. Science Robotics, 7(62):eabk2822, 2022
2022
-
[60]
G. B. Margolis and P. Agrawal. Walk these ways: Tuning robot control for generalization with multiplicity of behavior. In Conference on Robot Learning, pages 22–31. PMLR, 2023
2023
-
[61]
S. Choi, G. Ji, J. Park, H. Kim, J. Mun, J. H. Lee, and J. Hwangbo. Learning quadrupedal locomotion on deformable terrain. Science Robotics, 8(74):eade2256, 2023
2023
-
[62]
Caluwaerts, A
K. Caluwaerts, A. Iscen, J. C. Kew, W. Yu, T. Zhang, D. Freeman, K.-H. Lee, L. Lee, S. Sal- iceti, V . Zhuang, et al. Barkour: Benchmarking animal-level agility with quadruped robots. arXiv preprint arXiv:2305.14654, 2023
2023 arXiv
-
[63]
Stasica, A
M. Stasica, A. Bick, N. Bohlinger, O. Mohseni, J. Fritzsche, C. Hübler, J. Peters, and A. Sey- farth. Bridge the gap: Enhancing quadruped locomotion with vertical ground perturba- tions. In Under review, 2025. URL https://www.ias.informatik.tu-darmstadt.de/ uploads/Team/NicoBo...
2025
-
[64]
Zhuang, Z
Z. Zhuang, Z. Fu, J. Wang, C. Atkeson, S. Schwertfeger, C. Finn, and H. Zhao. Robot parkour learning. In Conference on Robot Learning (CoRL), 2023
2023
-
[65]
Cheng, K
X. Cheng, K. Shi, A. Agarwal, and D. Pathak. Extreme parkour with legged robots. In RoboLetics: Workshop on Robot Learning in Athletics@ CoRL 2023, 2023
2023
-
[66]
Siekmann, K
J. Siekmann, K. Green, J. Warila, A. Fern, and J. Hurst. Blind bipedal stair traversal via sim-to-real reinforcement learning. In Robotics: Science and Systems, 2021
2021
-
[67]
Kumar, Z
A. Kumar, Z. Li, J. Zeng, D. Pathak, K. Sreenath, and J. Malik. Adapting rapid motor adap- tation for bipedal robots. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1161–1168. IEEE, 2022
2022
-
[68]
Radosavovic, T
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath. Real-world hu- manoid locomotion with reinforcement learning. arXiv:2303.03381, 2023
2023 arXiv
-
[69]
Q. Liao, B. Zhang, X. Huang, X. Huang, Z. Li, and K. Sreenath. Berkeley humanoid: A research platform for learning-based control. arXiv preprint arXiv:2407.21781, 2024. 16
2024 arXiv
-
[70]
Zhuang, S
Z. Zhuang, S. Yao, and H. Zhao. Humanoid parkour learning. arXiv preprint arXiv:2406.10759, 2024
2024 arXiv
-
[71]
Chane-Sane, J
E. Chane-Sane, J. Amigo, T. Flayols, L. Righetti, and N. Mansard. Soloparkour: Constrained reinforcement learning for visual locomotion from privileged experience. In8th Annual Con- ference on Robot Learning, 2024
2024
-
[72]
Kaufmann, L
E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V . Koltun, and D. Scaramuzza. Champion-level drone racing using deep reinforcement learning. Nature, 620(7976):982– 987, 2023
2023
-
[73]
Kumar, Z
A. Kumar, Z. Fu, D. Pathak, and J. Malik. Rma: Rapid motor adaptation for legged robots. Robotics: Science and Systems XVII, 2021
2021
-
[74]
Rudin, D
N. Rudin, D. Hoeller, P. Reist, and M. Hutter. Learning to walk in minutes using mas- sively parallel deep reinforcement learning. InConference on Robot Learning, pages 91–100. PMLR, 2022
2022
-
[75]
Margolis, G
G. Margolis, G. Yang, K. Paigwar, T. Chen, and P. Agrawal. Rapid locomotion via reinforce- ment learning. In Robotics: Science and Systems, 2022
2022
-
[76]
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel. Sim-to-real transfer of robotic control with dynamics randomization. In 2018 IEEE international conference on robotics and automation (ICRA), pages 3803–3810. IEEE, 2018
2018
-
[77]
Campanaro, S
L. Campanaro, S. Gangapurwala, W. Merkt, and I. Havoutis. Learning and deploy- ing robust locomotion policies with minimal dynamics randomization. arXiv preprint arXiv:2209.12878, 2022
2022 arXiv
-
[78]
Smith, I
L. Smith, I. Kostrikov, and S. Levine. A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning. arXiv preprint arXiv:2208.07860, 2022
2022 arXiv
-
[79]
Smith, Y
L. Smith, Y . Cao, and S. Levine. Grow your limits: Continuous improvement with real-world rl for robotic locomotion. arXiv preprint arXiv:2310.17634, 2023
2023 arXiv
-
[80]
J. Levy, T. Westenbroek, and D. Fridovich-Keil. Learning to walk from three minutes of real-world data with semi-structured dynamics models. In 8th Annual Conference on Robot Learning, 2024
2024
-
[81]
Bohlinger, J
N. Bohlinger, J. Kinzel, D. Palenicek, L. Antczak, and J. Peters. Gait in eight: Efficient on- robot learning for omnidirectional quadruped locomotion. arXiv preprint arXiv:2503.08375, 2025
2025 arXiv
-
[82]
Jenelten, J
F. Jenelten, J. He, F. Farshidian, and M. Hutter. Dtc: Deep tracking control–a unifying ap- proach to model-based planning and reinforcement-learning for versatile and robust locomo- tion. arXiv preprint arXiv:2309.15462, 2023
2023 arXiv
-
[83]
Kasaei, M
M. Kasaei, M. Abreu, N. Lau, A. Pereira, and L. P. Reis. A cpg-based agile and versatile locomotion framework using proximal symmetry loss. arXiv preprint arXiv:2103.00928 , 2021
2021 arXiv
-
[84]
A. Zhao, J. Xu, M. Konakovi ´c-Lukovi´c, J. Hughes, A. Spielberg, D. Rus, and W. Matusik. Robogrammar: graph grammar for terrain-optimized robot design. ACM Transactions on Graphics (TOG), 39(6):1–16, 2020
2020
-
[85]
Azakami, H
T. Azakami, H. Kera, and K. Kawamoto. Adversarial body shape search for legged robots. In 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages 682–687. IEEE, 2022. 17
2022
-
[86]
Rajani, K
C. Rajani, K. Arndt, D. Blanco-Mulero, K. S. Luck, and V . Kyrki. Co-imitation: learning design and behaviour by imitation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 6200–6208, 2023
2023
-
[87]
Hazard, N
C. Hazard, N. Pollard, and S. Coros. Automated design of robotic hands for in-hand manip- ulation tasks. International Journal of Humanoid Robotics, 17(01):1950029, 2020
2020
-
[88]
Gupta, L
A. Gupta, L. Fan, S. Ganguli, and L. Fei-Fei. Metamorph: Learning universal controllers with transformers. arXiv preprint arXiv:2203.11931, 2022
2022 arXiv
-
[89]
Patel and S
A. Patel and S. Song. Get-zero: Graph embodiment transformer for zero-shot embodiment generalization. arXiv preprint arXiv:2407.15002, 2024
2024 arXiv
-
[90]
Cheng, Y
X. Cheng, Y . Ji, J. Chen, R. Yang, G. Yang, and X. Wang. Expressive whole-body control for humanoid robots. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors, Robotics: Science and Systems XX, Delft, The Netherlands, July 15-19, 2024, 2024. doi:10.15607/RSS. 202...
2024 doi
-
[91]
Bjorck, F
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, Linxi, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. LLontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wan...
- [92]
-
[93]
Sferrazza, D
C. Sferrazza, D. Huang, X. Lin, Y . Lee, and P. Abbeel. Humanoidbench: Simulated humanoid benchmark for whole-body locomotion and manipulation. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors, Robotics: Science and Systems XX, Delft, The Netherlands, July 15-19...
2024 doi
-
[94]
H. Shi, W. Wang, S. Song, and C. K. Liu. Toddlerbot: Open-source ml-compatible humanoid platform for loco-manipulation, 2025. URL https://arxiv.org/abs/2502.00893
2025
-
[95]
M. Liu, Z. Chen, X. Cheng, Y . Ji, R. Qiu, R. Yang, and X. Wang. Visual whole-body control for legged loco-manipulation. In P. Agrawal, O. Kroemer, and W. Burgard, ed- itors, Conference on Robot Learning, 6-9 November 2024, Munich, Germany , volume 270 of Proceedings of Machin...
2024
-
[96]
R. Yang, M. Zhang, N. Hansen, H. Xu, and X. Wang. Learning vision-guided quadrupedal locomotion end-to-end with cross-modal transformers. InThe Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net,
2022
-
[97]
T. He, C. Zhang, W. Xiao, G. He, C. Liu, and G. Shi. Agile but safe: Learning collision- free high-speed legged locomotion. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors, Robotics: Science and Systems XX, Delft, The Netherlands, July 15-19, 2024 , 2024. doi:1...
2024 doi
-
[98]
G. B. Margolis, G. Yang, K. Paigwar, T. Chen, and P. Agrawal. Rapid locomotion via reinforcement learning. Int. J. Robotics Res. , 43(4):572–587, 2024. doi:10.1177/ 02783649231224053. URL https://doi.org/10.1177/02783649231224053. 18
2024 doi
-
[99]
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y . Bai, D. Hafner, S. Bohez, and V . Vanhoucke. Sim- to-real: Learning agile locomotion for quadruped robots. In H. Kress-Gazit, S. S. Srinivasa, T. Howard, and N. Atanasov, editors, Robotics: Science and Systems XIV , Carnegie Mellon U...
2018 doi
-
[100]
Zhang, Y
H. Zhang, Y . Liu, J. Zhao, J. Chen, and J. Yan. Development of a Bionic Hexapod Robot for Walking on Unstructured Terrain. Journal of Bionic Engineering , 11(2):176–187, June
-
[101]
Z. Zang, M. Kawawa-Beaudan, W. Yu, T. Zhang, and A. Zakhor. Perceptive Hexapod Legged Locomotion for Climbing Joist Environments. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 2738–2745, Detroit, MI, USA, Oct. 2023. IEEE. ISBN 978-1...
2023
-
[102]
T. Qu, D. Li, A. Zakhor, W. Yu, and T. Zhang. Versatile locomotion skills for hexapod robots. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 6885–6892, 2024. doi:10.1109/IROS58592.2024.10801714
2024
-
[103]
Ouyang, H
W. Ouyang, H. Chi, J. Pang, W. Liang, and Q. Ren. Adaptive locomotion control of a hexapod robot via bio-inspired learning. Frontiers Neurorobotics, 15:627157, 2021. doi:10.3389/ FNBOT.2021.627157. URL https://doi.org/10.3389/fnbot.2021.627157
2021
-
[104]
Azayev and K
T. Azayev and K. Zimmerman. Blind Hexapod Locomotion in Complex Terrain with Gait Adaptation Using Deep Reinforcement Learning and Classification.J Intell Robot Syst, 2020
2020
-
[105]
Chiu, Y .-C
J.-R. Chiu, Y .-C. Huang, H.-C. Chen, K.-Y . Tseng, and P.-C. Lin. Development of a Running Hexapod Robot with Differentiated Front and Hind Leg Morphology and Functionality. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 3710–3717, ...
2020
-
[106]
A CPG-based locomo- tion control architecture for hexapod robot
Haitao Yu, Wei Guo, Jing Deng, Mantian Li, and Hegao Cai. A CPG-based locomo- tion control architecture for hexapod robot. In 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages 5615–5621, Tokyo, Nov. 2013. IEEE. doi: 10.1109/iros.2013.6697170. URL...
2013
-
[107]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6000–6010, Red Hook, NY , USA,
-
[108]
Mittal, C
M. Mittal, C. Yu, Q. Yu, J. Liu, N. Rudin, D. Hoeller, J. L. Yuan, R. Singh, Y . Guo, H. Mazhar, A. Mandlekar, B. Babich, G. State, M. Hutter, and A. Garg. Orbit: A unified simulation framework for interactive robot learning environments. IEEE Robotics and Automation Let- ters...
2023
-
[109]
van der Maaten and G
L. van der Maaten and G. Hinton. Visualizing data using t-sne. Journal of Ma- chine Learning Research, 9(86):2579–2605, 2008. URL http://jmlr.org/papers/v9/ vandermaaten08a.html
2008
-
[110]
Jolliffe
I. Jolliffe. Principal component analysis. Springer Verlag, New York, 2002
2002
-
[111]
McInnes and J
L. McInnes and J. Healy. UMAP: uniform manifold approximation and projection for dimen- sion reduction. CoRR, abs/1802.03426, 2018. URL http://arxiv.org/abs/1802.03426. 19
2018 arXiv
-
[112]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7
2019
-
[113]
starting
I. Loshchilov and F. Hutter. SGDR: stochastic gradient descent with warm restarts. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings . OpenReview.net, 2017. URL https: //openreview.net/forum?...
2017
-
[2014]
doi:10.1016/S1672-6529(14)60041-X
ISSN 2543-2141. doi:10.1016/S1672-6529(14)60041-X. URL https://doi.org/ 10.1016/S1672-6529(14)60041-X
-
[2017]
ISBN 9781510860964
Curran Associates Inc. ISBN 9781510860964
-
[2020]
URL https://arxiv.org/abs/2001.08361
2001 arXiv
-
[2021]
doi:10.18653/v1/2021.acl-long.295
Association for Computational Linguistics. doi:10.18653/v1/2021.acl-long.295. URL https://aclanthology.org/2021.acl-long.295/
2021 doi
-
[2022]
URL https://openreview.net/forum?id=nhnJ3oo6AB
-
[2023]
URL https://openreview.net/forum?id=vuSI9mhDaBZ
- [2024]
-
[2713]
URL https://proceedings.mlr.press/v270/kim25c.html
PMLR, 2024. URL https://proceedings.mlr.press/v270/kim25c.html
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.