Pith. sign in

REVIEW 3 major objections 3 minor 5 cited by

Design and Control of a Bipedal Robotic Character

T0 review · 3 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A puppeteered bipedal robot can act out expressive, unscripted live shows, and the authors report roughly ten hours of public runtime with no falls.

desk verdict A credible and well-integrated system paper for entertainment robotics; the evaluation is thin in places but the central engineering claim holds. read the letter →

arxiv 2501.05204 v1 pith:6OSPYSXR submitted 2025-01-09 cs.RO cs.LG

classification cs.ROcs.LG
keywords bipedalrobotcharactercharacter-drivenmechanicaldesignreinforcementlearninglocomotionimitationfromanimationsim-to-realtransferpuppeteeringinterfaceengineentertainmentrobotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that expressive, artist-directed motion and dynamic balance can be unified in a single legged-robot system by co-designing the robot body with its motion repertoire, training separate reinforcement-learning policies to imitate artist-authored motions, and letting an animation engine blend those policies behind a two-joystick interface. A human operator can then puppeteer the robot through live, unscripted performances, switching between standing, walking, and one-shot emotional animations. The authors built a small bipedal character with creatively driven proportions and show functions such as actuated antennas and illuminated eyes, and deployed up to three robots in public settings, accumulating about ten hours of runtime without a single fall. If the workflow holds, custom expressive robot characters for entertainment and human engagement can be developed quickly without giving up dynamic mobility.

What carries the argument

The key object is the path frame, a moving coordinate frame that anchors every artist-authored motion to the world and makes transitions between policies consistent. Each motion reference is stored in path coordinates and mapped to world coordinates by a generator; during standing the frame converges to the feet, during walking it integrates commanded velocities, and for episodic motions its trajectory is part of the artistic input. The path frame, together with a phase signal fed to the policy as a feature vector, lets separate policies be switched and blended seamlessly because every commanded pose is expressed in a common, robot-relative frame. The other load-bearing mechanism is the imitation-reward reinforcement-learning stack: a weighted sum of pose, velocity, joint, and contact rewards compares simulated motion to kinematic references, while actuator models and domain randomization bridge the gap between simulation and hardware.

What would settle it

Run the physical robot through a fixed script of the same walking speeds, triggered animations, and walk-to-stand transitions used in the paper on a hard floor, and compare its falls, joint tracking errors, and foot-contact timing with the same script executed in simulation under identical commands. If the real tracking errors and fall rate exceed the simulated distribution substantially, or if the phase-synchronized policy switch becomes visibly discontinuous on hardware, the sim-to-real assumption is falsified.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a single workflow—character-driven mechanical design, policies trained to imitate artist-authored motion references, and an animation engine that composes background, triggered, and joystick layers—can put expressive, dynamic performance on a physical bipedal robot and keep it believable in front of an audience. The design separates motion into three temporal types: perpetual standing, periodic walking driven by a phase signal and velocity commands, and episodic one-shot animations. Each type gets its own reinforcement-learning policy, and the animation engine's low-dimensional command signals, together with phase-aware switching, make the transitions look continuous to an outside observer. The evidence offered includes joint-tracking error tables, torque-limit checks during an episodic jump, responses to pushes and small obstacles, and roughly ten hours of public runtime across three robots with no falls.

Load-bearing premise

Everything rests on the assumption that policies trained only in simulation transfer to the physical robot without fine-tuning; if the simulator's contact, actuator, or mass models are not faithful enough, the live performance could fail despite the reported ten hours.

Editorial extensions

If this is right

  • A character robot can be developed in under a year from off-the-shelf actuators and 3D-printed structural parts, because the mechanical design is tuned to creative intent rather than extreme performance.
  • An operator without a robotics background can puppeteer expressive performances after training, because the interface separates gaze from posture and routes all commands through the animation engine.
  • Separate policies for standing, walking, and episodic motions can be swapped on the fly without visible discontinuities, as long as transitions are made phase-aware.
  • The same workflow can produce expressive characters with non-anthropomorphic morphologies, because the motion-reference format and path-frame interface are not tied to a human body plan.
  • Public deployments with up to three robots operating simultaneously accumulated about ten hours of runtime with zero falls.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that the path-frame plus phase-signal command interface could serve as a standard API between animation tools and reinforcement-learning robot controllers, letting artists author for a robot without touching the control stack.
  • My extension would be a perceptual study measuring whether audiences rate a puppeteered character as more alive than the same motions played autonomously; the reported bystander questions about whether the robot can see suggest this is directly testable.
  • The ten-hour no-fall record is an operational result rather than a controlled benchmark, so a systematic stress test with scripted pushes and varied floor surfaces would separate policy robustness from operator skill.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents a complete workflow for creating a bipedal robot character: mechanical design driven by artistic intent, simulation-based training of multiple reinforcement-learning policies for standing, walking, and episodic motions, a runtime animation engine that blends background, triggered, and joystick-driven animation layers, and a two-joystick puppeteering interface. The system is evaluated through joint-tracking errors (Table III), velocity-following plots (Fig. 7), torque-limit analysis during a jump (Fig. 8), policy-transition plots (Fig. 9), qualitative video comparisons against alternative RL formulations, and roughly ten hours of public deployment without a fall.

Significance. If the results hold, the paper provides a valuable demonstration that a character-driven, non-anthropomorphic biped can execute expressive, artist-directed motions while retaining robust dynamic mobility. The strengths include a clearly described divide-and-conquer policy architecture, detailed actuator system identification and domain randomization parameters, and a rare public-deployment record that substantiates the central claim. The paper is also notable for explicitly separating perpetual, periodic, and episodic motion types and for integrating show functions into the animation pipeline. The provided appendices give enough quantitative detail to make the approach reproducible in principle.

major comments (3)
  1. [Section VII-A, Table III] The mean absolute tracking errors in Table III are reported without any indication of the number of trials, the variance across runs, or the specific command and phase conditions under which they were measured. Since these numbers are the central quantitative support for the tracking claim, please report the number of episodes, the mean and standard deviation (or range), and the experimental protocol used to generate the data.
  2. [Section VII-B] The comparison with alternative RL formulations is entirely qualitative, relying on statements such as "as shown in the video" and "visually identical motion." If the paper claims any comparative advantage, please provide quantitative metrics for the alternative policies (e.g., tracking error, foot clearance, joint acceleration, or energy consumption); otherwise, explicitly frame the comparison as qualitative so readers can calibrate the strength of the claim.
  3. [Section V-E / Appendix B / Fig. 8] The sim-to-real argument rests on actuator models identified on an isolated test bench and randomized only within isolated-actuator parameter ranges. The paper does not address in-situ effects such as shared-battery voltage sag, thermal drift, or coupled power draw during sustained high-power episodes, even though the presented jump reaches the actuator torque limits. Please add either measurements of battery voltage and motor temperature during representative shows or a discussion of the expected operating margin relative to the randomized envelope, so that the robustness claim is not overstated.
minor comments (3)
  1. [Section VII-A, Fig. 7] The statement that "the robot is responsive and follows all commands closely" would be stronger with a quantitative measure of velocity tracking error or latency, rather than only the visual comparison in the figure.
  2. [Section V-A / Table I] The hat notation for target state quantities (e.g., \hat{p}, \hat{q}) is used throughout Table I but is not explicitly defined in the text; please add a short explanation near Eq. (1).
  3. [General] The manuscript contains the watermark "Disney Confidential - Do Not Distribute" on every page; this is unusual for a journal submission and should be removed before final submission, or its presence must be explained in the cover letter.

Circularity Check

0 steps flagged · score 1.0 of 10

No material circularity: the paper is an empirical robotic systems result whose central claims rest on hardware deployment, RL training, and independent component evaluations, not on a derivation that re-imports its own assumptions.

full rationale

The paper's claims are validated experimentally rather than derived from a closed analytical chain. The RL policies are trained with imitation rewards against artist-authored kinematic references defined in Sec. V-A and Tab. I, and the paper then evaluates tracking error on the physical robot (Tab. III, Figs. 7-8); the references are inputs, not outputs of the control stack. The actuator model in App. B is identified on a test bench and used both in simulation and to compute torque limits in Fig. 8; this is system identification feeding a simulator, with the comparison to measured torques serving as validation, not as a fitted 'prediction' of the paper's main claim. The procedural gait generator cited as [14] is an author-group prior tool used as an animation-input stage, but the central claim is the integrated RL control and runtime system, which is demonstrated against external deployments (roughly 10 h without a fall) and compared with alternative RL formulations in Sec. VII-B. No load-bearing step reduces by construction to its own inputs; the self-citation is incidental and not doing justificatory work for the paper's conclusions.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on simulation fidelity, hand-tuned rewards and disturbance schedules, and the assumption that the reference motion generation pipeline yields feasible motions. No new theoretical entities are introduced.

free parameters (3)
  • Reward weights in Tab. I = e.g., leg joint position 15.0, neck joint position 100.0, contact 1.0, survival 20.0
    Hand-tuned to balance imitation, regularization, and survival; no sensitivity analysis or systematic tuning is reported, yet they directly shape the learned behavior.
  • Episodic reward weight schedule (Eq. 13) = w_extra and interval [phi_start, phi_end] for 'excited motion' and 'jump'
    Ad hoc adjustment added to avoid the policy 'cheating' (keeping toes on the ground) and to emphasize key motion aspects; chosen by hand.
  • Domain randomization ranges (Tab. V) = Force 0-150 N, torque 0-15 N m, durations as listed
    Disturbance curriculum and parameter ranges selected by the authors; no evidence matches real-world disturbances, but they are needed for robust sim-to-real transfer.
assumptions (5)
  • domain assumption Rigid-body dynamics in Isaac Gym accurately model the physical robot (Sec. V-E).
    The entire training and sim-to-real transfer relies on the fidelity of the simulation, including mass, contacts, and actuator dynamics.
  • domain assumption Actuator models identified on a test bench transfer to the assembled robot (App. B).
    The actuator model parameters are measured on single actuators and assumed to hold for the assembled system with randomized offsets.
  • domain assumption Artist-authored reference motions are dynamically feasible after processing with model-based tools (Sec. V-A).
    The walking references are generated with MPC and inverse dynamics, and episodic references are assumed to be imitable by the RL policy.
  • standard math PPO with the listed hyperparameters converges to policies that generalize under domain randomization (App. A).
    Standard RL convergence assumptions; the paper provides no proof of convergence, only empirical results.
  • domain assumption Path frame and phase parameterization is sufficient to represent and transfer motions (Sec. V-A).
    The path frame and phase signal are the core mechanism for transitions and command conditioning; if this representation is insufficient, the policy blending would fail.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Design and Control of a Bipedal Robotic Character." pith.science (2026). https://pith.science/paper/6OSPYSXR

@misc{pith2026250105204,
  author       = {Pith},
  title        = {Pith review of: Design and Control of a Bipedal Robotic Character},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6OSPYSXR}},
  note         = {Machine review of arXiv:2501.05204}
}
read the original abstract

Legged robots have achieved impressive feats in dynamic locomotion in challenging unstructured terrain. However, in entertainment applications, the design and control of these robots face additional challenges in appealing to human audiences. This work aims to unify expressive, artist-directed motions and robust dynamic mobility for legged robots. To this end, we introduce a new bipedal robot, designed with a focus on character-driven mechanical features. We present a reinforcement learning-based control architecture to robustly execute artistic motions conditioned on command signals. During runtime, these command signals are generated by an animation engine which composes and blends between multiple animation sources. Finally, an intuitive operator interface enables real-time show performances with the robot. The complete system results in a believable robotic character, and paves the way for enhanced human-robot engagement in various contexts, in entertainment robotics and beyond.

Figures

Figures reproduced from arXiv: 2501.05204 by the authors.

Figure 1
Figure 1. Three instances of our robotic character performing an unscripted [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Our character design and control pipeline consists of animation, mechatronic design, reinforcement learning, and run-time tools. Animation and [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Mechanical design of our robotic character. The robot has 5 degrees [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Path frame illustrations for standing (top-left), and a top-view during walking (bottom-left). During standing, the path frame converges towards [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: A selection of joystick commands during standing. Posture control [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Commanded path velocities (dashed) and measured torso velocities [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Measured joint torques (solid) and velocity-dependent torque limits [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Policy actions across policy transitions during a short motion [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]
Figure 10
Figure 10. Figure 10: (Top) The operator uses the proposed puppeteering interface to act out a scene where the robot discovers a paper roll and decides to kick it under [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]
Figure 11
Figure 11. Figure 11: Joystick mapping during standing for the torso and head yaw offset [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Steam Deck [43] layout. TABLE VII PUPPETEERING BUTTON MAPPING Button Effect Menu Trigger a safety mode called motion stop. This forces a transition to standing and freezes the joint setpoints with high po￾sition gains after waiting 0.5 s. View Slowly move all joints t…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning

    cs.RO 2025-05 conditional novelty 7.0 of 10

    SPI-Active identifies legged-robot physical parameters via massive parallel sampling and uses Fisher-information-optimal command sequences to collect informative real-world data, improving sim-to-real transfer on quad...

  2. Learning to Walk in Costume: Adversarial Motion Priors for Aesthetically Constrained Humanoids

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A reinforcement learning controller with human-motion imitation priors produced stable standing and walking on Cosmo, a top-heavy, vision-less entertainment humanoid with shell-restricted joints.

  3. Demonstrating Berkeley Humanoid Lite: An Open-source, Accessible, and Customizable 3D-printed Humanoid Robot

    cs.RO 2025-04 conditional novelty 6.0 of 10

    A low-cost, open-source humanoid platform using 3D-printed cycloidal actuators is demonstrated with reinforcement-learning locomotion and teleoperated manipulation.

  4. Robust RL Control for Bipedal Locomotion with Closed Kinematic Chains

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A reinforcement-learning gait controller that explicitly models closed kinematic chains outperforms one trained on a simplified serial model, both in simulation and on the physical TopA robot.

  5. System Identification of Thrust and Torque Characteristics for a Bipedal Robot with Integrated Propulsion

    cs.RO 2025-04 conditional novelty 3.0 of 10

    The paper fits a thrust equation with two empirical constants and measures the KV380 motor torque constant as 0.0311 Nm/A, but the thrust validation depends on fitted parameters and a coarse 200-node CFD.

Reference graph

Works this paper leans on

52 extracted references · 40 canonical work pages · cited by 5 Pith papers

  1. [1]

    How Boston Dynamics taught its robots to dance

    Evan Ackerman. How Boston Dynamics taught its robots to dance. IEEE Spectrum , Jan. 7 2021. URL https://spectrum.ieee.org/ how-boston-dynamics-taught-its-robots-to-dance

  2. [2]

    Maya, 2023

    Autodesk, INC. Maya, 2023. URL https://autodesk.com/ maya

  3. [3]

    Real-time dance generation to music for a legged robot

    Thomas Bi, P ´eter Fankhauser, Dario Bellicoso, and Marco Hutter. Real-time dance generation to music for a legged robot. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2018

  4. [4]

    Offline motion libraries and online mpc for advanced mobility skills

    Marko Bjelonic, Ruben Grandia, Moritz Geilinger, Oliver Harley, Vivian S Medeiros, Vuk Pajovic, Edo Jelavic, Stelian Coros, and Marco Hutter. Offline motion libraries and online mpc for advanced mobility skills. The International Journal of Robotics Research , 41(9-10): 903–924, 2022

  5. [5]

    Adversarial motion priors make good substitutes for complex reward functions

    Alejandro Escontrela, Xue Bin Peng, Wenhao Yu, Tingnan Zhang, Atil Iscen, Ken Goldberg, and Pieter Abbeel. Adversarial motion priors make good substitutes for complex reward functions. In 2022 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 25–32. IEEE, 2022

  6. [6]

    Iterative Cauchy Thresholding: Regularisation with a heavy-tailed prior

    M. Fujita, Y . Kuroki, T. Ishida, and T.T. Doi. Au- tonomous behavior control architecture of entertain- ment humanoid robot SDR-4X. In Proceedings 2003 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2003) (Cat. No.03CH37453) , vol- ume 1, pages 960–967 vol.1, 2003. doi: 10.1109/IROS. 2003.1250752

  7. [7]

    Aibo: Toward the era of digital crea- tures

    Masahiro Fujita. Aibo: Toward the era of digital crea- tures. The International Journal of Robotics Research , 20(10):781–794, 2001

  8. [8]

    An open architec- ture for robot entertainment

    Masahiro Fujita and Koji Kageyama. An open architec- ture for robot entertainment. In Proceedings of the first international conference on Autonomous agents , pages 435–442, 1997

Show all 52 references
  1. [9]

    Mecha- tronic design of NAO humanoid

    David Gouaillier, Vincent Hugel, Pierre Blazevic, Chris Kilner, J ´erˆome Monceaux, Pascal Lafourcade, Brice Marnier, Julien Serre, and Bruno Maisonnier. Mecha- tronic design of NAO humanoid. In 2009 IEEE inter- national conference on robotics and automation , pages 769–774. I...

  2. [10]

    Doc: Differentiable optimal control for retargeting motions onto legged robots

    Ruben Grandia, Farbod Farshidian, Espen Knoop, Chris- tian Schumacher, Marco Hutter, and Moritz B¨acher. Doc: Differentiable optimal control for retargeting motions onto legged robots. ACM Trans. Graph., 42(4), jul 2023. ISSN 0730-0301

  3. [11]

    Contact-aided invariant extended kalman filtering for robot state estimation

    Ross Hartley, Maani Ghaffari, Ryan M Eustice, and Jessy W Grizzle. Contact-aided invariant extended kalman filtering for robot state estimation. The Inter- national Journal of Robotics Research , 39(4):402–430,

  4. [12]

    CoMic: Complementary task learning & mimicry for reusable skills

    Leonard Hasenclever, Fabio Pardo, Raia Hadsell, Nicolas Heess, and Josh Merel. CoMic: Complementary task learning & mimicry for reusable skills. In Hal Daum ´e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning , volume 119 of Pr...

  5. [13]

    Anymal parkour: Learning agile navigation for quadrupedal robots

    David Hoeller, Nikita Rudin, Dhionis Sako, and Marco Hutter. Anymal parkour: Learning agile navigation for quadrupedal robots. Science Robotics , 9(88):eadi7566, 2024

  6. [14]

    Hopkins, Georg Wiedebach, Kyle Cesare, Jared Bishop, Espen Knoop, and Moritz B ¨acher

    Michael A. Hopkins, Georg Wiedebach, Kyle Cesare, Jared Bishop, Espen Knoop, and Moritz B ¨acher. In- teractive design of stylized walking gaits for robotic characters. ACM Transactions On Graphics (TOG), 2024

  7. [15]

    Walk this way: To be useful around people, robots need to learn how to move like we do

    Jonathan Hurst. Walk this way: To be useful around people, robots need to learn how to move like we do. IEEE Spectrum , 56(3):30–51, 2019. doi: 10.1109/ MSPEC.2019.8651932

  8. [16]

    Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger

    Marco Hutter, Christian Gehring, Dominic Jud, Andreas Lauber, C. Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger. ANYmal - a highly mo- bile and dynamic quadrup...

  9. [17]

    Learning agile and dynamic motor skills for legged robots

    Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, and Marco Hutter. Learning agile and dynamic motor skills for legged robots. Science Robotics, 4(26):eaau5872, 2019

  10. [18]

    Robovie: an interactive humanoid robot

    Hiroshi Ishiguro, Tetsuo Ono, Michita Imai, Takeshi Maeda, Takayuki Kanda, and Ryohei Nakatsu. Robovie: an interactive humanoid robot. Industrial robot: An international journal, 28(6):498–504, 2001

  11. [19]

    Rl + model-based con- trol: Using on-demand optimal control to learn versatile legged locomotion

    Dongho Kang, Jin Cheng, Miguel Zamora, Fatemeh Zargarbashi, and Stelian Coros. Rl + model-based con- trol: Using on-demand optimal control to learn versatile legged locomotion. IEEE Robotics and Automation Letters, 8(10):6619–6626, 2023. doi: 10.1109/LRA.2023. 3307008

  12. [20]

    Emys—emotive head of a social robot

    Jan Kedzierski, Robert Muszy ´nski, Carsten Zoll, Adam Oleksy, and Mirela Frontkiewicz. Emys—emotive head of a social robot. International Journal of Social Robotics, 5:237–249, 2013

  13. [21]

    Design of a momentum-based control framework and application to the humanoid robot atlas

    Twan Koolen, Sylvain Bertrand, Gray Thomas, Tomas De Boer, Tingfan Wu, Jesper Smith, Johannes Engls- berger, and Jerry Pratt. Design of a momentum-based control framework and application to the humanoid robot atlas. International Journal of Humanoid Robotics , 13 (01):1650007, 2016

  14. [22]

    Evoking agency: attention model and behavior control in a robotic art installation

    Christian Kroos, Damith C Herath, and Stelarc. Evoking agency: attention model and behavior control in a robotic art installation. Leonardo, 45(5):401–407, 2012

  15. [23]

    Data-driven biped control

    Yoonsang Lee, Sungeun Kim, and Jehee Lee. Data-driven biped control. In ACM SIGGRAPH 2010 Papers , SIG- GRAPH ’10, New York, NY , USA, 2010. Association for Computing Machinery. ISBN 9781450302104

  16. [24]

    Animated cassie: A dynamic relatable robotic character

    Zhongyu Li, Christine Cummings, and Koushil Sreenath. Animated cassie: A dynamic relatable robotic character. In 2020 IEEE/RSJ International Conference on Intel- ligent Robots and Systems (IROS) , pages 3739–3746. IEEE, 2020

  17. [25]

    Reinforcement learning for robust parameterized loco- motion control of bipedal robots

    Zhongyu Li, Xuxin Cheng, Xue Bin Peng, Pieter Abbeel, Sergey Levine, Glen Berseth, and Koushil Sreenath. Reinforcement learning for robust parameterized loco- motion control of bipedal robots. In 2021 IEEE Interna- tional Conference on Robotics and Automation (ICRA) , pages 28...

  18. [26]

    Isaac gym: High performance gpu based physics simulation for robot learning

    Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, and Gavriel State. Isaac gym: High performance gpu based physics simulation for robot learning. In J. Van- schoren and S. Yeu...

  19. [27]

    The icub humanoid robot: An open-systems platform for research in cognitive development

    Giorgio Metta, Lorenzo Natale, Francesco Nori, Giulio Sandini, David Vernon, Luciano Fadiga, Claes V on Hof- sten, Kerstin Rosander, Manuel Lopes, Jos ´e Santos- Victor, et al. The icub humanoid robot: An open-systems platform for research in cognitive development. Neural netw...

  20. [28]

    Learning robust perceptive locomotion for quadrupedal robots in the wild

    Takahiro Miki, Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, and Marco Hutter. Learning robust perceptive locomotion for quadrupedal robots in the wild. Science Robotics , 7(62), 2022. doi: 10.1126/ scirobotics.abk2822

  21. [29]

    Computer animation of human walking: a survey

    Franck Multon, Laure France, Marie-Paule Cani- Gascuel, and Giles Debunne. Computer animation of human walking: a survey. The journal of visualization and computer animation , 10(1):39–54, 1999

  22. [30]

    Intuitive and flexible user interface for creating whole body motions of biped humanoid robots

    Shin’ichiro Nakaoka, Shuuji Kajita, and Kazuhito Yokoi. Intuitive and flexible user interface for creating whole body motions of biped humanoid robots. In 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems, pages 1675–1682. IEEE, 2010

  23. [31]

    A mass- produced sociable humanoid robot: Pepper: The first machine of its kind

    Amit Kumar Pandey and Rodolphe Gelin. A mass- produced sociable humanoid robot: Pepper: The first machine of its kind. IEEE Robotics & Automation Magazine, 25(3):40–48, 2018. doi: 10.1109/MRA.2018. 2833157

  24. [32]

    Deepmimic: Example-guided deep re- inforcement learning of physics-based character skills

    Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne. Deepmimic: Example-guided deep re- inforcement learning of physics-based character skills. ACM Transactions On Graphics (TOG) , 37(4):1–14, 2018

  25. [33]

    Learning agile robotic locomotion skills by imitating animals

    Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang- Wei Lee, Jie Tan, and Sergey Levine. Learning agile robotic locomotion skills by imitating animals. Robotics: Science and Systems , 2020

  26. [34]

    Amp: Adversarial motion priors for stylized physics-based character control

    Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Transac- tions on Graphics (ToG) , 40(4):1–20, 2021

  27. [35]

    Choregraphe: a graphical tool for hu- manoid robot programming

    Emmanuel Pot, J ´erˆome Monceaux, Rodolphe Gelin, and Bruno Maisonnier. Choregraphe: a graphical tool for hu- manoid robot programming. In RO-MAN 2009-The 18th IEEE International Symposium on Robot and Human Interactive Communication, pages 46–51. IEEE, 2009

  28. [36]

    The psychosocial effects of a companion robot: a randomized controlled trial

    Hayley Robinson, Bruce MacDonald, Ngaire Kerse, and Elizabeth Broadbent. The psychosocial effects of a companion robot: a randomized controlled trial. Journal of the American Medical Directors Association , 14(9): 661–667, 2013

  29. [37]

    Learning to walk in minutes using massively parallel deep reinforcement learning

    Nikita Rudin, David Hoeller, Philipp Reist, and Marco Hutter. Learning to walk in minutes using massively parallel deep reinforcement learning. In Conference on Robot Learning, pages 91–100. PMLR, 2022

  30. [38]

    Proximal policy optimization algorithms

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017

  31. [39]

    Sim-to-real learning of all common bipedal gaits via periodic reward composition

    Jonah Siekmann, Yesh Godse, Alan Fern, and Jonathan Hurst. Sim-to-real learning of all common bipedal gaits via periodic reward composition. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 7309–7315. IEEE, 2021

  32. [40]

    Blind bipedal stair traversal via sim- to-real reinforcement learning

    Jonah Siekmann, Kevin Green, John Warila, Alan Fern, and Jonathan Hurst. Blind bipedal stair traversal via sim- to-real reinforcement learning. Robotics: Science and Systems XVII, 2021

  33. [41]

    On human motion imitation by humanoid robot

    Wael Suleiman, Eiichi Yoshida, Fumio Kanehiro, Jean- Paul Laumond, and Andr ´e Monin. On human motion imitation by humanoid robot. In 2008 IEEE International conference on robotics and automation , pages 2697–

  34. [42]

    Sim-to-real: Learning agile locomotion for quadruped robots

    Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, and Vincent Vanhoucke. Sim-to-real: Learning agile locomotion for quadruped robots. In Proceedings of Robotics: Science and Systems , Pittsburgh, Pennsylvania, June 2018. doi: 10.15607...

  35. [43]

    Steam deck, 2023

    Valve Corporation. Steam deck, 2023. URL https://www. steamdeck.com/

  36. [44]

    van Breemen

    Albert J.N. van Breemen. Animation engine for believ- able interactive user-interface robots. In 2004 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems (IROS)(IEEE Cat. No. 04CH37566) , volume 3, pages 2873–2878. IEEE, 2004

  37. [45]

    van Breemen

    Albert J.N. van Breemen. Bringing robots to life: Ap- plying principles of animation to robots. In Proceedings of Shapping Human-Robot Interaction workshop held at CHI, volume 2004, pages 143–144. Citeseer, 2004

  38. [46]

    Do robot performance and behavioral style affect human trust? a multi-method approach

    Rik van den Brule, Ron Dotsch, Gijsbert Bijlstra, Daniel HJ Wigboldus, and Pim Haselager. Do robot performance and behavioral style affect human trust? a multi-method approach. International journal of social robotics, 6:519–531, 2014

  39. [47]

    Robot expressive motions: A survey of generation and evaluation meth- ods

    Gentiane Venture and Dana Kuli ´c. Robot expressive motions: A survey of generation and evaluation meth- ods. J. Hum.-Robot Interact. , 8(4), nov 2019. doi: 10.1145/3344286

  40. [48]

    Survey on human–robot collaboration in indus- trial settings: Safety, intuitive interfaces and applications

    Valeria Villani, Fabio Pini, Francesco Leali, and Cristian Secchi. Survey on human–robot collaboration in indus- trial settings: Safety, intuitive interfaces and applications. Mechatronics, 55:248–266, 2018. ISSN 0957-4158. doi: https://doi.org/10.1016/j.mechatronics.2018.02.009

  41. [49]

    Wensing, Albert Wang, Sangok Seok, David Otten, Jeffrey Lang, and Sangbae Kim

    Patrick M. Wensing, Albert Wang, Sangok Seok, David Otten, Jeffrey Lang, and Sangbae Kim. Proprioceptive actuator design in the mit cheetah: Impact mitigation and high-bandwidth physical interaction for dynamic legged robots. IEEE Transactions on Robotics , 33(3):509–522, 2017

  42. [50]

    Optimization-based control for dynamic legged robots

    Patrick M Wensing, Michael Posa, Yue Hu, Adrien Escande, Nicolas Mansard, and Andrea Del Prete. Optimization-based control for dynamic legged robots. IEEE Transactions on Robotics , 2023

  43. [51]

    excited motion

    Pierre-brice Wieber. Trajectory free linear model predic- tive control for stable walking in the presence of strong perturbations. In 2006 6th IEEE-RAS International Conference on Humanoid Robots , pages 137–142, 2006. doi: 10.1109/ICHR.2006.321375. APPENDIX A POLICY TRAINING ...

  44. [2020]

    doi: 10.1177/0278364919894385

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.