Pith. sign in

REVIEW 4 major objections 5 minor 58 references

Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Robot-Trains-Robot claims a force-sensing robot arm can teach a humanoid to walk and to swing up, using only minutes of real-world training.

desk verdict RTR is a credible and useful hardware system, but the headline speed-doubling claim is not backed by a direct speed measurement and the walking reward loop is coupled to the teacher's own force/treadmill feedback, so the quantitative claims need careful handling. read the letter →

arxiv 2508.12252 v2 pith:RX6MALDH submitted 2025-08-17 cs.RO

classification cs.RO
keywords humanoidrobotsreal-worldreinforcementlearningsim-to-realtransferdynamicslatentoptimizationteacher-studentrobotcompliantarmwalkingspeedtrackingswing-up
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Robot-Trains-Robot (RTR), a system in which a robot arm with force sensing acts as teacher for a humanoid student: it supports the student safely, supplies reward signals that would otherwise be unavailable, runs an automatic curriculum, detects failures, and resets the robot without human help. The central claim is that this setup makes real-world reinforcement learning practical on humanoids. In the walking task, fine-tuning a single dynamics-encoded latent variable with RTR doubles the zero-shot walking speed after about 20 minutes of real-world training. In the swing-up task, the humanoid learns a periodic swing-up motion from scratch within about 15 minutes of real-world interaction. If correct, this is evidence that the missing ingredient for real-world humanoid learning is not the RL algorithm alone but a physical teacher that closes the safety, reward, and reset loops.

What carries the argument

The load-bearing object is the dynamics latent $\mathbf{z}$, a single vector that encodes environment physics. In Stage 1, an encoder $f_\phi$ maps randomized physics parameters $\mu^{(i)}$ to $\mathbf{z}^{(i)}$, and Feature-wise Linear Modulation (FiLM) layers modulate the actor's hidden states by scaling and shifting them, so the policy becomes dynamics-conditioned. Stage 2 optimizes a universal latent $\tilde{\mathbf{z}}$ shared across all simulated environments to give a reliable start. Stage 3 freezes the actor and FiLM parameters and fine-tunes $\tilde{\mathbf{z}}$ in the real world with PPO, so real-world adaptation is reduced to optimizing one low-dimensional latent instead of the full network. The hardware counterpart is the admittance-controlled teacher arm with a force-torque sensor, plus an optional treadmill whose speed is PD-controlled from tether force and torso pitch; these provide safety, the speed proxy used as reward, and automatic reset.

What would settle it

Measure the humanoid's true forward velocity independently with motion capture or with an overground trial while RTR trains; if the treadmill speed stays high while true walking speed does not rise, or if the learned policy fails to transfer to a treadmill without the tether feedback loop, the central speed-doubling claim would be falsified.

Watch

Extended reading notes

Core claim

The paper's central discovery is that a teacher arm with force-torque sensing can supply everything a humanoid policy needs to learn in the real world: safe exploration via compliant support, reward via measured interaction, curriculum via scheduled support height and helping or perturbing arm motions, and automation via failure detection and resets. For sim-to-real adaptation, the paper proposes a three-stage pipeline: a dynamics-aware policy is pretrained in domain-randomized simulation with a latent vector encoded from physics parameters and injected through FiLM layers; a universal latent is optimized across all simulated environments; then only that latent is fine-tuned in the real world while the actor and FiLM layers stay frozen. This one-latent fine-tuning is what yields the reported speed-tracking improvement. For learning from scratch, the same physical infrastructure provides a helping and perturbing arm schedule and offline critic pretraining, letting the humanoid discover a swing-up motion directly in the real world.

Load-bearing premise

The walking result assumes that the treadmill speed, which is itself computed from the humanoid's pull on the elastic tether and its torso pitch, faithfully measures how fast the humanoid is actually walking; if the policy instead learns to exploit the tether to raise its own reward, the speed-tracking result would not be about walking speed.

Editorial extensions

If this is right

  • If the central claim holds, real-world fine-tuning of a humanoid walking policy can be reduced to optimizing a single latent vector, making adaptation far more data-efficient than full-network fine-tuning.
  • Real-world learning from scratch becomes feasible for tasks that are difficult to simulate, such as cable-suspended swing-up, because the teacher arm provides safety and a physical curriculum without a scripted policy.
  • The same infrastructure can double as a data-collection, failure-detection, and reset system, cutting human supervision to the point of long, mostly unattended training sessions.
  • The system is platform-agnostic in principle: any robot arm or crane with force sensing could serve as teacher for larger humanoids whenever payload and workspace allow.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the walking reward loop couples treadmill speed to tether force, so a policy that learns to pull the tether could raise its own reward without walking faster; an overground transfer test would separate true walking from tether exploitation.
  • Editorial inference: the helping and perturbing arm schedule is a physical curriculum that pumps energy into the system, suggesting the same mechanism could train other underactuated or unstable robots whenever an external agent can safely add or remove energy.
  • Editorial inference: if the universal-latent initialization generalizes, the same three-stage recipe could serve as a sim-to-real warm start for other legged robots, with the latent encoding terrain, friction, or payload rather than only humanoid dynamics.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Robot-Trains-Robot (RTR), a system in which a force-sensing robot arm acts as a teacher that supports, guides, rewards, perturbs, resets, and schedules a small humanoid student during real-world reinforcement learning. The learning method has three stages: training a dynamics-conditioned policy with FiLM modulation across domain-randomized simulation environments, optimizing a universal dynamics latent in simulation, and then fine-tuning only that latent in the real world with PPO. The paper reports two hardware demonstrations on the ToddlerBot platform: fine-tuning a walking policy for treadmill speed tracking, claimed to double the zero-shot walking speed with 20 minutes of real-world training, and learning a swing-up behavior from scratch in 15 minutes. Ablations cover arm compliance, arm height scheduling, latent versus policy fine-tuning, FiLM learning rate, and a comparison with RMA.

Significance. If the quantitative claims are supportable, RTR is a significant systems contribution: it is one of the few demonstrations of autonomous real-world humanoid RL with minimal human intervention, and the dynamics-latent fine-tuning pipeline is a sensible way to keep real-world updates low-dimensional. Strengths include three-seed ablations for the main comparisons, explicit baselines (fine-tuning the base policy, fine-tuning a residual policy, fixed-arm and fixed-latent variants, and RMA), a FiLM learning-rate ablation with zero-latent evaluation, and the use of an open-source, low-cost humanoid platform. The main unresolved issue is measurement validity: the headline results are evaluated through the teacher's own force and treadmill feedback loop rather than through independent measurements of the learned skill.

major comments (4)
  1. [Section 3.2, Eq. (9), and Section 4.1] The headline claim that RTR 'doubles the zero-shot walking speed' is not backed by an independent measurement of walking speed. The reward in Eq. (3) is computed from v, which the paper says is 'approximated' by the treadmill speed, and Appendix C.3 defines that treadmill speed as v = v_base + k1^p Fx + k2^p psi, where Fx is the force the humanoid exerts on the elastic tether and psi is torso pitch. A policy can therefore increase its own reward by pulling on the tether or adopting a posture that drives psi without increasing actual gait speed. No before-and-after speed value, speed-tracking error curve, or maximum-speed number appears in Section 4.1 or Table 1; Table 1 reports only stability metrics at a single belt speed of 0.15 m/s. Please provide an independent measurement of the robot's gait speed (for example, motion-capture torso velocity, foot-contact-based speed, or camera tracking) and report the actual pre/post speed values that support the 'doubles' claim.
  2. [Section 3.3, Eq. (4), and Figure 5] The swing-up result is also measured through the teacher's own sensor loop. The reward in Eq. (4) is the FFT amplitude of the force sensor at the dominant frequency, while the teacher arm's helping motion is phase-aligned to the same estimated swing phase, with x_t = x0 + A_arm cos(theta_t). The paper reports no independent rope-angle, IMU-based tilt, or vision-based swing-height measurement, so the learned behavior could be an interaction with the arm rather than a true pendulum swing-up. Please add an independent swing-angle or swing-height trace, or otherwise validate quantitatively that the force amplitude tracks the actual swing amplitude.
  3. [Appendix C.5] Appendix C.5, labeled 'Real-world Learning Details', is an empty heading with no content. This is exactly the section that would document the real-world protocol for both tasks, including reset procedures, batch timing, evaluation protocol, and the measurements behind the 20-minute and 15-minute claims. Its absence makes the headline experiments difficult to reproduce and prevents the reader from assessing the proxy-reward concerns above. The section should be filled in or its content should be integrated into the main experimental details.
  4. [Figure 4 and Section 4.1] The walking ablations are evaluated with the same proxy reward that is being optimized. The y-axis 'linear velocity tracking rewards' is computed from the Eq. (9) treadmill speed, so the training curves conflate genuine policy improvement with exploitation of the force and pitch feedback loop. The 'better data efficiency' conclusion should be re-stated in terms of the measured walking outcome once an independent speed measurement is added; as written, the conclusion is about proxy reward rather than gait speed.
minor comments (5)
  1. [Appendix D] The heading 'Rapad Motor Adaptation' should be 'Rapid Motor Adaptation'.
  2. [Appendix C.2] The text refers to 'appliance control'; from the context this appears to be a typo for 'admittance control'.
  3. [Section 4.2] The sentence 'we apply an larger entropy coefficient' should be 'we apply a larger entropy coefficient'.
  4. [Section 3.3] The expression 'theta_0 = 30 degrees' uses an unusual degree symbol; use the standard notation for consistency.
  5. [References] Reference [10], a market research report, is an unusual source for the claim about industrial arm payload capacity; a technical datasheet or manufacturer specification would be more appropriate.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the walking and swing-up claims are real-world RL measurements with ablations; the treadmill/force proxy is a measurement-validity caveat, not a derivation tautology.

full rationale

Walking the derivation chain, the paper does not derive a headline result from equations that already contain that result. The three-stage latent optimization is a standard PPO/FiLM pipeline: a dynamics latent is optimized in simulation and then fine-tuned in the real world, and the claimed gains are empirical learning curves, not algebraic consequences of the reward definition. The walking reward uses the treadmill speed as a proxy for robot velocity (Eq. 3, with v approximated by the treadmill speed), and Eq. 9 defines that treadmill speed as a PD output depending on tether force and torso pitch; in principle, a policy could inflate its own reward by pulling on the tether. Similarly, the swing-up reward (Eq. 4) is the FFT amplitude of the force sensor that the teacher arm can directly excite during helping. These are serious measurement-channel/proxy concerns that belong to correctness and reproducibility review, and the paper's Section 6 explicitly acknowledges the lack of ground-reaction force sensing and proposes a force plate as future work. However, they are not circular in the required sense: no fitted parameter is renamed as a prediction, no result is mathematically forced by self-citation, and no uniqueness theorem is imported from the authors' prior work. Self-citations ([9], [35], [55]) provide the open-source hardware, encoder architecture, and three-stage training template; they are implementation support rather than load-bearing evidence for the central claim. The empty Appendix C.5 removes protocol detail and weakens verification, but missing detail is not circularity. The balance of evidence—ablations in Figure 4, FiLM ablations in Appendix B, and RMA comparison in Appendix D—gives the central claims independent empirical content. Score 2 reflects minor self-citations and the proxy-reward caveat, not a circular derivation chain.

Assumptions & free parameters 7 free parameters · 4 assumptions · 1 invented entities

The central claims rest on hand-chosen constants (reward shapes, arm schedule, treadmill gains), the small-angle derivation of the swing-up target, the proxy assumption that treadmill speed equals walking speed, and the fidelity of the randomized simulator. No new physical entities are postulated; the dynamics latent is a standard learnable conditioning variable whose role is supported by the paper's own ablations.

free parameters (7)
  • walking reward shaping sigma = 100
    Exponent in Eq. 3 (r = exp(-sigma (v - v_target)^2)); hand-chosen, strongly shapes the reward landscape.
  • swing-up reward shaping alpha and target amplitude A_target = alpha = 0.005, A_target approx mg*theta0 with m = 3.5 kg, theta0 = 30 deg
    Eq. 4; theta0 is a chosen expected amplitude, and m differs slightly from the 3.4 kg stated for ToddlerBot in Section 3.1.
  • arm guidance/perturbation amplitude A_arm = 0.05 m
    Position target amplitude for helping and perturbing modes in the swing-up task (Section 3.3).
  • arm Z schedule = linear decrease of 0.02 m over 5e4 env steps
    Curriculum schedule that reduces support to near zero (Section 3.2).
  • treadmill PD gains = k1^p = 0.2, k2^p = -5, v_base = 0.1 m/s, speed cap 0.24 m/s
    Eq. 9 and Appendix C.3; these constants determine the treadmill speed that is used as the robot velocity estimate and reward signal.
  • FiLM layer learning rate = 5e-5
    Selected from the sweep in Appendix B by performance on the ablation; thus tuned on the reported ablation data.
  • dynamics latent z (universal and real-world) = 1024-dim vector; values not reported
    Stage 2 optimizes z-tilde over randomized sim environments (Eq. 2) and Stage 3 fine-tunes z* against the real-world reward; values are learned and not reported, so the real-world result cannot be reproduced without code.
assumptions (4)
  • domain assumption Small-angle pendulum approximation for the swing-up target amplitude A_target approx m*g*theta0
    Eq. 4 states A_target is 'derived under the small-angle approximation of a pendulum swing'; if the physical swing is not small-angle, the target amplitude may not correspond to the intended swing height.
  • domain assumption Treadmill speed proxies the humanoid's true walking velocity
    Section 3.2 approximates v with treadmill speed; Appendix C.3 defines treadmill speed as a PD function of tether force Fx and torso pitch (Eq. 9), coupling the measurement and the reward.
  • domain assumption Elastic-rope tether permits free planar motion with smooth force transmission
    Section 3.1 claims rope elasticity is crucial for smoother force transmission and avoiding abrupt forces; the learning results depend on the tether not constraining or biasing the policy.
  • domain assumption The domain-randomized simulator is a valid training ground for the walking policy
    Standard sim-to-real premise (Section 3.2, Appendix C.1); simulator fidelity and randomization ranges are taken from [9] and [55] without independent validation in this paper.
invented entities (1)
  • dynamics latent z (FiLM-conditioned environment code) independent evidence
    purpose: Encodes environment physics from domain-randomized simulation parameters and modulates the policy via FiLM layers for sim-to-real adaptation; the universal and real-world variants are optimized as initialization and adaptation parameters.
    This is a standard learnable latent variable from context-based meta-RL (RMA [30], Yu et al. [43]), not a new physical entity. The paper's own FiLM ablation (Appendix B) shows evaluation reward collapses with a zero latent when the FiLM learning rate is large, providing in-paper evidence it carries information.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids." pith.science (2026). https://pith.science/paper/RX6MALDH

@misc{pith2026250812252,
  author       = {Pith},
  title        = {Pith review of: Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RX6MALDH}},
  note         = {Machine review of arXiv:2508.12252}
}
read the original abstract

Simulation-based reinforcement learning (RL) has significantly advanced humanoid locomotion tasks, yet direct real-world RL from scratch or adapting from pretrained policies remains rare, limiting the full potential of humanoid robots. Real-world learning, despite being crucial for overcoming the sim-to-real gap, faces substantial challenges related to safety, reward design, and learning efficiency. To address these limitations, we propose Robot-Trains-Robot (RTR), a novel framework where a robotic arm teacher actively supports and guides a humanoid robot student. The RTR system provides protection, learning schedule, reward, perturbation, failure detection, and automatic resets. It enables efficient long-term real-world humanoid training with minimal human intervention. Furthermore, we propose a novel RL pipeline that facilitates and stabilizes sim-to-real transfer by optimizing a single dynamics-encoded latent variable in the real world. We validate our method through two challenging real-world humanoid tasks: fine-tuning a walking policy for precise speed tracking and learning a humanoid swing-up task from scratch, illustrating the promising capabilities of real-world humanoid learning realized by RTR-style systems. See https://robot-trains-robot.github.io/ for more info.

Figures

Figures reproduced from arXiv: 2508.12252 by the authors.

Figure 1
Figure 1. Robot Trains Robot (RTR). We pro￾pose RTR for automatic real-world policy adapta￾tion and learning with a robot arm as the teacher and a humanoid robot as the student. Recent advances in training reinforcement learning (RL) policies using massive parallel simulation environments have yielded remark￾able results in humanoid locomotion tasks [1, 2, 3, 4, 5]. These methods demonstrate the abil￾ity to deploy on physical… view at source ↗
Figure 2
Figure 2. System Setup. We illustrate the system architecture and component interactions. The system consists of two groups: robot teachers and robot students. The teachers include a robot arm with an F/T sensor, a mini PC, and an optional treadmill for locomotion tasks; the students include a humanoid robot and a workstation for policy training. The four types of lines represent physical interaction, data transmission, contr… view at source ↗
Figure 3
Figure 3. Sim-to-real Fine-tuning Algorithm. We illustrate our sim-to-real finetuning process. First, we train a dynamics-aware policy in simulation via domain randomization (DR), encoding en￾vironment physics into a latent vector. Next, we optimize a universal latent across diverse simulation environments to initialize real-world training. Finally, we refine the latent and train a new critic in the real world. Orange denotes… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Walking Ablation. This experiment aims to evaluate the effectiveness of arm feedback control and latent vector finetuning. We present the linear velocity tracking rewards during training and evaluation, with the arm schedule shown at the bottom center. All variants are…
Figure 5
Figure 5. Figure 5: Swing-up Ablation. We illustrate the swing-up setup and experiment results. (a) The humanoid is suspended from a robot arm and uses its legs to build momentum and maximize rope angle. (b) We compare helping and perturbing arm schedules against a fixed-arm baseline and …
Figure 6
Figure 6. Figure 6: FiLM learning rate ablation. This experiment aims to evaluate the effect of FiLM layer learning rates. All experiments are run under seven seeds. Vertical bars indicate standard deviation [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 41 canonical work pages

  1. [1]

    T. He, J. Gao, W. Xiao, Y . Zhang, Z. Wang, J. Wang, Z. Luo, G. He, N. Sobanbab, C. Pan, Z. Yi, G. Qu, K. Kitani, J. Hodgins, L. J. Fan, Y . Zhu, C. Liu, and G. Shi. ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills, Feb. 2025

  2. [2]

    T. He, W. Xiao, T. Lin, Z. Luo, Z. Xu, Z. Jiang, J. Kautz, C. Liu, G. Shi, X. Wang, L. Fan, and Y . Zhu. HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots, Mar. 2025

  3. [3]

    Radosavovic, T

    I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath. Real-world humanoid locomotion with reinforcement learning. Science Robotics, 9(89):eadi9579, Apr. 2024. doi: 10.1126/scirobotics.adi9579

  4. [4]

    Radosavovic, S

    I. Radosavovic, S. Kamat, T. Darrell, and J. Malik. Learning Humanoid Locomotion over Challenging Terrain, Oct. 2024

  5. [5]

    Zhuang, S

    Z. Zhuang, S. Yao, and H. Zhao. Humanoid Parkour Learning. In 8th Annual Conference on Robot Learning, Sept. 2024

  6. [6]

    Tobin, R

    J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 23–30, Sept. 2017. doi: 10.1109/IROS.2017.8202133

  7. [7]

    Perez, F

    E. Perez, F. Strub, H. de Vries, V . Dumoulin, and A. Courville. Film: Visual reasoning with a general conditioning layer, 2017. URL https://arxiv.org/abs/1709.07871

  8. [8]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms, 2017. URL https://arxiv.org/abs/1707.06347

Show all 58 references
  1. [9]

    H. Shi, W. Wang, S. Song, and C. K. Liu. ToddlerBot: Open-Source ML-Compatible Hu- manoid Platform for Loco-Manipulation, Feb. 2025

  2. [10]

    Industrial robotic arm market size, share, growth trends, regional share, competitive intelligence, forecast report 2025–2037, March 2025

    Research Nester. Industrial robotic arm market size, share, growth trends, regional share, competitive intelligence, forecast report 2025–2037, March 2025. URL https://www. researchnester.com/reports/industrial-robotic-arm-market/6763 . Accessed: 2025-04-28

  3. [11]

    X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel. Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 3803–3810, May 2018. doi: 10.1109/ICRA.2018.8460528

  4. [12]

    Z. Yuan, T. Wei, S. Cheng, G. Zhang, Y . Chen, and H. Xu. Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning, Oct. 2024

  5. [13]

    H. Ha, Y . Gao, Z. Fu, J. Tan, and S. Song. UMI-on-Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers. In 8th Annual Conference on Robot Learning, Sept. 2024

  6. [14]

    Rudin, D

    N. Rudin, D. Hoeller, P. Reist, and M. Hutter. Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning. In Proceedings of the 5th Conference on Robot Learn- ing, pages 91–100. PMLR, Jan. 2022

  7. [15]

    Cheng, K

    X. Cheng, K. Shi, A. Agarwal, and D. Pathak. Extreme Parkour with Legged Robots. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 11443–11450, May

  8. [16]

    J. Tan, T. Zhang, E. Coumans, A. Iscen, Y . Bai, D. Hafner, S. Bohez, and V . Vanhoucke. Sim-to-Real: Learning Agile Locomotion For Quadruped Robots. In Robotics: Science and Systems XIV, volume 14, June 2018. ISBN 978-0-9923747-4-7. 10

  9. [17]

    F. Shi, Y . Kojio, T. Makabe, T. Anzai, K. Kojima, K. Okada, and M. Inaba. Reference- free learning bipedal motor skills via assistive force curricula. In A. Billard, T. Asfour, and O. Khatib, editors, Robotics Research, pages 304–320, Cham, 2023. Springer Nature Switzer- land...

  10. [18]

    Akkaya, M

    OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang. Solving Rubik’s Cube with a Robot Hand, Oct. 2019

  11. [19]

    T. Chen, M. Tippur, S. Wu, V . Kumar, E. Adelson, and P. Agrawal. Visual dexterity: In-hand reorientation of novel and complex object shapes. Science Robotics , 8(84):eadc9244, Nov

  12. [20]

    Y . Chen, C. Wang, L. Fei-Fei, and K. Liu. Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation. In 7th Annual Conference on Robot Learning , Aug. 2023

  13. [21]

    Y . Chen, C. Wang, Y . Yang, and K. Liu. Object-Centric Dexterous Manipulation from Human Motion Data. In 8th Annual Conference on Robot Learning , Sept. 2024

  14. [22]

    T. Lin, K. Sachdev, L. Fan, J. Malik, and Y . Zhu. Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids, Feb. 2025

  15. [23]

    Z. Fu, Q. Zhao, Q. Wu, G. Wetzstein, and C. Finn. HumanPlus: Humanoid Shadowing and Imitation from Humans. In 8th Annual Conference on Robot Learning , Sept. 2024

  16. [24]

    Gu, Y .-J

    X. Gu, Y .-J. Wang, X. Zhu, C. Shi, Y . Guo, Y . Liu, and J. Chen. Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning. In Robotics: Science and Systems XX . Robotics: Science and Systems Foundation, July 2024. ISBN 9798990284807. ...

  17. [25]

    Haarnoja, B

    T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, J. Humplik, M. Wulfmeier, S. Tun- yasuvunakool, N. Y . Siegel, R. Hafner, M. Bloesch, K. Hartikainen, A. Byravan, L. Hasen- clever, Y . Tassa, F. Sadeghi, N. Batchelor, F. Casarini, S. Saliceti, C. Game, N. Sreendra, K...

  18. [26]

    Chebotar, A

    Y . Chebotar, A. Handa, V . Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox. Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience. In 2019 International Conference on Robotics and Automation (ICRA) , pages 8973–8979, May

  19. [27]

    A. Z. Ren, H. Dai, B. Burchfiel, and A. Majumdar. AdaptSim: Task-Driven Simulation Adap- tation for Sim-to-Real Transfer. In Proceedings of The 7th Conference on Robot Learning , pages 3434–3452. PMLR, Dec. 2023

  20. [28]

    Huang, X

    P. Huang, X. Zhang, Z. Cao, S. Liu, M. Xu, W. Ding, J. Francis, B. Chen, and D. Zhao. What Went Wrong? Closing the Sim-to-Real Gap via Differentiable Causal Discovery. In Proceedings of The 7th Conference on Robot Learning , pages 734–760. PMLR, Dec. 2023

  21. [29]

    Schoettler, A

    G. Schoettler, A. Nair, J. A. Ojea, S. Levine, and E. Solowjow. Meta-Reinforcement Learning for Robotic Industrial Insertion Tasks. In 2020 IEEE/RSJ International Conference on Intel- ligent Robots and Systems (IROS) , pages 9728–9735, Oct. 2020. doi: 10.1109/IROS45743. 2020.9340848

  22. [30]

    Kumar, Z

    A. Kumar, Z. Fu, D. Pathak, and J. Malik. RMA: Rapid Motor Adaptation for Legged Robots, July 2021. 11

  23. [31]

    H. Qi, A. Kumar, R. Calandra, Y . Ma, and J. Malik. In-Hand Object Rotation via Rapid Motor Adaptation. In 6th Annual Conference on Robot Learning , Aug. 2022

  24. [32]

    Kumar, Z

    A. Kumar, Z. Li, J. Zeng, D. Pathak, K. Sreenath, and J. Malik. Adapting Rapid Motor Adap- tation for Bipedal Robots. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1161–1168, Oct. 2022. doi: 10.1109/IROS47612.2022.9981091

  25. [33]

    Y . Sun, W. L. Ubellacker, W.-L. Ma, X. Zhang, C. Wang, N. V . Csomay-Shanklin, M. Tomizuka, K. Sreenath, and A. D. Ames. Online Learning of Unknown Dynamics for Model-Based Controllers in Legged Locomotion. IEEE Robotics and Automation Letters , 6 (4):8442–8449, Oct. 2021. IS...

  26. [34]

    Zhang, L

    Y . Zhang, L. Ke, A. Deshpande, A. Gupta, and S. Srinivasa. Cherry-Picking with Reinforce- ment Learning. In Robotics: Science and Systems XIX , volume 19, July 2023. ISBN 978-0- 9923747-9-2

  27. [35]

    K. Lei, Z. He, C. Lu, K. Hu, Y . Gao, and H. Xu. Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization. InThe Twelfth International Conference on Learning Representations, Oct. 2023

  28. [36]

    Xiong, R

    H. Xiong, R. Mendonca, K. Shaw, and D. Pathak. Adaptive Mobile Manipulation for Articu- lated Objects In the Open World, Jan. 2024

  29. [37]

    Jiang, C

    Y . Jiang, C. Wang, R. Zhang, J. Wu, and L. Fei-Fei. TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction. In 8th Annual Conference on Robot Learning , Sept. 2024

  30. [38]

    J. Luo, C. Xu, J. Wu, and S. Levine. Precise and Dexterous Robotic Manipulation via Human- in-the-Loop Reinforcement Learning, Mar. 2025

  31. [39]

    Gupta, R

    A. Gupta, R. Mendonca, Y . Liu, P. Abbeel, and S. Levine. Meta-reinforcement learning of structured exploration strategies, 2018. URL https://arxiv.org/abs/1802.07245

  32. [40]

    Rakelly, A

    K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine. Efficient off-policy meta- reinforcement learning via probabilistic context variables, 2019. URL https://arxiv.org/ abs/1903.08254

  33. [41]

    W. Yu, J. Tan, Y . Bai, E. Coumans, and S. Ha. Learning fast adaptation with meta strategy optimization, 2020. URL https://arxiv.org/abs/1909.12995

  34. [42]

    X. B. Peng, E. Coumans, T. Zhang, T.-W. Lee, J. Tan, and S. Levine. Learning agile robotic locomotion skills by imitating animals, 2020. URL https://arxiv.org/abs/2004.00784

  35. [43]

    W. Yu, V . C. Kumar, G. Turk, and C. K. Liu. Sim-to-real transfer for biped locomotion, 2019. URL https://arxiv.org/abs/1903.01390

  36. [44]

    Huang, Z

    S. Huang, Z. Zhang, T. Liang, Y . Xu, Z. Kou, C. Lu, G. Xu, Z. Xue, and H. Xu. MEN- TOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning, Oct. 2024

  37. [45]

    K. Xu, Z. Hu, R. Doshi, A. Rovinsky, V . Kumar, A. Gupta, and S. Levine. Dexterous Ma- nipulation from Images: Autonomous Real-World RL via Substep Guidance. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 5938–5945, May 2023. doi: 10.1109/ICRA4...

  38. [46]

    Mendonca, E

    R. Mendonca, E. Panov, B. Bucher, J. Wang, and D. Pathak. Continuously Improving Mobile Manipulation with Autonomous Real-World RL. In Proceedings of The 8th Conference on Robot Learning, pages 5204–5219. PMLR, Jan. 2025. 12

  39. [47]

    S. Ha, P. Xu, Z. Tan, S. Levine, and J. Tan. Learning to Walk in the Real World with Minimal Human Effort. In Proceedings of the 2020 Conference on Robot Learning , pages 1110–1120. PMLR, Oct. 2021

  40. [48]

    Smith, I

    L. Smith, I. Kostrikov, and S. Levine. A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning, Aug. 2022

  41. [49]

    P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg. DayDreamer: World Models for Physical Robot Learning. In Proceedings of The 6th Conference on Robot Learning , pages 2226–2240. PMLR, Mar. 2023

  42. [50]

    Smith, Y

    L. Smith, Y . Cao, and S. Levine. Grow your limits: Continuous improvement with real-world rl for robotic locomotion, 2023. URL https://arxiv.org/abs/2310.17634

  43. [51]

    Bloesch, J

    M. Bloesch, J. Humplik, V . Patraucean, R. Hafner, T. Haarnoja, A. Byravan, N. Y . Siegel, S. Tunyasuvunakool, F. Casarini, N. Batchelor, F. Romano, S. Saliceti, M. Riedmiller, S. M. A. Eslami, and N. Heess. Towards real robot learning in the wild: A case study in bipedal loco...

  44. [52]

    ROBOTIS OP3 e-Manual

    ROBOTIS. ROBOTIS OP3 e-Manual. https://emanual.robotis.com/docs/en/ platform/op3/introduction/, 2024. Accessed: 2025-07-29

  45. [53]

    Maples and J

    J. Maples and J. Becker. Experiments in force control of robotic manipulators. In 1986 IEEE International Conference on Robotics and Automation Proceedings, volume 3, pages 695–702, Apr. 1986. doi: 10.1109/ROBOT.1986.1087590

  46. [54]

    Y . Hou, Z. Liu, C. Chi, E. Cousineau, N. Kuppuswamy, S. Feng, B. Burchfiel, and S. Song. Adaptive Compliance Policy: Learning Approximate Compliance for Diffusion Guided Con- trol, Mar. 2025

  47. [55]

    Ground Truth

    Y . Liang, T. Xu, K. Hu, G. Jiang, F. Huang, and H. Xu. Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , Nov. 2024. 13 A State and Action Representation A...

  48. [2019]

    doi: 10.1109/ICRA.2019.8793789

  49. [2023]

    doi: 10.1126/scirobotics.adc9244

  50. [2024]

    doi: 10.1109/ICRA57147.2024.10610200

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.