Pith. sign in

REVIEW 4 cited by

Real-World Humanoid Locomotion with Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.03381 v2 pith:VHLLZOWD submitted 2023-03-06 cs.RO cs.LG

classification cs.ROcs.LG
keywords humanoidadaptenvironmentscontrollerhistorylearninglocomotionmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humanoid robots that can autonomously operate in diverse environments have the potential to help address labour shortages in factories, assist elderly at homes, and colonize new planets. While classical controllers for humanoid robots have shown impressive results in a number of settings, they are challenging to generalize and adapt to new environments. Here, we present a fully learning-based approach for real-world humanoid locomotion. Our controller is a causal transformer that takes the history of proprioceptive observations and actions as input and predicts the next action. We hypothesize that the observation-action history contains useful information about the world that a powerful transformer model can use to adapt its behavior in-context, without updating its weights. We train our model with large-scale model-free reinforcement learning on an ensemble of randomized environments in simulation and deploy it to the real world zero-shot. Our controller can walk over various outdoor terrains, is robust to external disturbances, and can adapt in context.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. Shared Control of Holonomic Wheelchairs through Reinforcement Learning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    An RL policy trained in Isaac Gym and tested in Gazebo and on a real DAA V1 wheelchair translates 2D joystick commands into collision-free 3D motion for a holonomic wheelchair.

  3. Generalized Locomotion in Out-of-distribution Conditions with Robust Transformer

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A transformer with body tokenization and consistent dropout generalizes to unseen leg damages and sensor noise while trained on limited dynamics and clean observations.

  4. EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A state-conditioned executable motion prior network modifies upper-body motion targets so a humanoid can imitate human gestures while maintaining balance.

Pith tools