Pith. sign in

REVIEW 8 cited by

Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05695 v2 pith:YHM4IMBA submitted 2024-04-08 cs.RO cs.AIcs.LGcs.SYeess.SY

classification cs.ROcs.AIcs.LGcs.SYeess.SY
keywords humanoidhumanoid-gymframeworkrobottransferzero-shotenvironmentisaac
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humanoid-Gym is an easy-to-use reinforcement learning (RL) framework based on Nvidia Isaac Gym, designed to train locomotion skills for humanoid robots, emphasizing zero-shot transfer from simulation to the real-world environment. Humanoid-Gym also integrates a sim-to-sim framework from Isaac Gym to Mujoco that allows users to verify the trained policies in different physical simulations to ensure the robustness and generalization of the policies. This framework is verified by RobotEra's XBot-S (1.2-meter tall humanoid robot) and XBot-L (1.65-meter tall humanoid robot) in a real-world environment with zero-shot sim-to-real transfer. The project website and source code can be found at: https://sites.google.com/view/humanoid-gym/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Handroid: Bridging Dexterous Hand and Humanoid

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A single 27-DoF body doubles as an anthropomorphic dexterous hand and a 0.33 m desktop humanoid, with a unified control stack for manipulation, locomotion, and embodiment switching.

  2. KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A robot control method that adaptively tightens motion-tracking reward tolerances achieves lower tracking errors on dynamic skills and transfers zero-shot to a real humanoid.

  3. RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control

    cs.RO 2025-06 reject novelty 6.0 of 10

    RLPF uses reinforcement learning with a physics-simulator tracking reward and an alignment verification module to fine-tune a large text-to-motion model for physically feasible humanoid motions.

  4. H2-COMPACT: Human-Humanoid Co-Manipulation via Adaptive Contact Trajectory Policies

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A hierarchical framework that maps wrist force/torque into velocity commands and then into stable leg motions lets a humanoid robot carry loads cooperatively with a human using only haptic cues.

  5. PRISM: Polynomial Representations for Interaction-Structured Motor Control

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Explicit low-degree factorized polynomial proprioceptive features improve robot RL and imitation policies beyond matched-capacity MLPs and induce sensorless compliance-like contact behavior in simulation.

  6. Learning Roller-Skating Motions of Humanoid Robots Based on Adversarial Motion Priors

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Independent AMP-PPO pipelines from retargeted mocap learn Pump Glide and Push Glide on a passive-wheel humanoid, with simulation metrics and real-robot trials.

  7. GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    GBC unifies MoCap retargeting and imitation learning into one framework that trains whole-body humanoid policies across multiple robot morphologies in simulation.

  8. Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A neural behavior classifier trained on teleoperated joint trajectories is proposed and tested as an offline meta-evaluator for imitation learning policies in human-robot interaction.

Pith tools