REVIEW 4 major objections 5 minor 58 references
Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Robot-Trains-Robot claims a force-sensing robot arm can teach a humanoid to walk and to swing up, using only minutes of real-world training.
desk verdict RTR is a credible and useful hardware system, but the headline speed-doubling claim is not backed by a direct speed measurement and the walking reward loop is coupled to the teacher's own force/treadmill feedback, so the quantitative claims need careful handling. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dynamics latent $\mathbf{z}$, a single vector that encodes environment physics. In Stage 1, an encoder $f_\phi$ maps randomized physics parameters $\mu^{(i)}$ to $\mathbf{z}^{(i)}$, and Feature-wise Linear Modulation (FiLM) layers modulate the actor's hidden states by scaling and shifting them, so the policy becomes dynamics-conditioned. Stage 2 optimizes a universal latent $\tilde{\mathbf{z}}$ shared across all simulated environments to give a reliable start. Stage 3 freezes the actor and FiLM parameters and fine-tunes $\tilde{\mathbf{z}}$ in the real world with PPO, so real-world adaptation is reduced to optimizing one low-dimensional latent instead of the full network. The hardware counterpart is the admittance-controlled teacher arm with a force-torque sensor, plus an optional treadmill whose speed is PD-controlled from tether force and torso pitch; these provide safety, the speed proxy used as reward, and automatic reset.
What would settle it
Measure the humanoid's true forward velocity independently with motion capture or with an overground trial while RTR trains; if the treadmill speed stays high while true walking speed does not rise, or if the learned policy fails to transfer to a treadmill without the tether feedback loop, the central speed-doubling claim would be falsified.
Extended reading notes
Core claim
The paper's central discovery is that a teacher arm with force-torque sensing can supply everything a humanoid policy needs to learn in the real world: safe exploration via compliant support, reward via measured interaction, curriculum via scheduled support height and helping or perturbing arm motions, and automation via failure detection and resets. For sim-to-real adaptation, the paper proposes a three-stage pipeline: a dynamics-aware policy is pretrained in domain-randomized simulation with a latent vector encoded from physics parameters and injected through FiLM layers; a universal latent is optimized across all simulated environments; then only that latent is fine-tuned in the real world while the actor and FiLM layers stay frozen. This one-latent fine-tuning is what yields the reported speed-tracking improvement. For learning from scratch, the same physical infrastructure provides a helping and perturbing arm schedule and offline critic pretraining, letting the humanoid discover a swing-up motion directly in the real world.
Load-bearing premise
The walking result assumes that the treadmill speed, which is itself computed from the humanoid's pull on the elastic tether and its torso pitch, faithfully measures how fast the humanoid is actually walking; if the policy instead learns to exploit the tether to raise its own reward, the speed-tracking result would not be about walking speed.
Editorial extensions
If this is right
- If the central claim holds, real-world fine-tuning of a humanoid walking policy can be reduced to optimizing a single latent vector, making adaptation far more data-efficient than full-network fine-tuning.
- Real-world learning from scratch becomes feasible for tasks that are difficult to simulate, such as cable-suspended swing-up, because the teacher arm provides safety and a physical curriculum without a scripted policy.
- The same infrastructure can double as a data-collection, failure-detection, and reset system, cutting human supervision to the point of long, mostly unattended training sessions.
- The system is platform-agnostic in principle: any robot arm or crane with force sensing could serve as teacher for larger humanoids whenever payload and workspace allow.
Reading between the lines
- Editorial inference: the walking reward loop couples treadmill speed to tether force, so a policy that learns to pull the tether could raise its own reward without walking faster; an overground transfer test would separate true walking from tether exploitation.
- Editorial inference: the helping and perturbing arm schedule is a physical curriculum that pumps energy into the system, suggesting the same mechanism could train other underactuated or unstable robots whenever an external agent can safely add or remove energy.
- Editorial inference: if the universal-latent initialization generalizes, the same three-stage recipe could serve as a sim-to-real warm start for other legged robots, with the latent encoding terrain, friction, or payload rather than only humanoid dynamics.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Robot-Trains-Robot (RTR), a system in which a force-sensing robot arm acts as a teacher that supports, guides, rewards, perturbs, resets, and schedules a small humanoid student during real-world reinforcement learning. The learning method has three stages: training a dynamics-conditioned policy with FiLM modulation across domain-randomized simulation environments, optimizing a universal dynamics latent in simulation, and then fine-tuning only that latent in the real world with PPO. The paper reports two hardware demonstrations on the ToddlerBot platform: fine-tuning a walking policy for treadmill speed tracking, claimed to double the zero-shot walking speed with 20 minutes of real-world training, and learning a swing-up behavior from scratch in 15 minutes. Ablations cover arm compliance, arm height scheduling, latent versus policy fine-tuning, FiLM learning rate, and a comparison with RMA.
Significance. If the quantitative claims are supportable, RTR is a significant systems contribution: it is one of the few demonstrations of autonomous real-world humanoid RL with minimal human intervention, and the dynamics-latent fine-tuning pipeline is a sensible way to keep real-world updates low-dimensional. Strengths include three-seed ablations for the main comparisons, explicit baselines (fine-tuning the base policy, fine-tuning a residual policy, fixed-arm and fixed-latent variants, and RMA), a FiLM learning-rate ablation with zero-latent evaluation, and the use of an open-source, low-cost humanoid platform. The main unresolved issue is measurement validity: the headline results are evaluated through the teacher's own force and treadmill feedback loop rather than through independent measurements of the learned skill.
major comments (4)
- [Section 3.2, Eq. (9), and Section 4.1] The headline claim that RTR 'doubles the zero-shot walking speed' is not backed by an independent measurement of walking speed. The reward in Eq. (3) is computed from v, which the paper says is 'approximated' by the treadmill speed, and Appendix C.3 defines that treadmill speed as v = v_base + k1^p Fx + k2^p psi, where Fx is the force the humanoid exerts on the elastic tether and psi is torso pitch. A policy can therefore increase its own reward by pulling on the tether or adopting a posture that drives psi without increasing actual gait speed. No before-and-after speed value, speed-tracking error curve, or maximum-speed number appears in Section 4.1 or Table 1; Table 1 reports only stability metrics at a single belt speed of 0.15 m/s. Please provide an independent measurement of the robot's gait speed (for example, motion-capture torso velocity, foot-contact-based speed, or camera tracking) and report the actual pre/post speed values that support the 'doubles' claim.
- [Section 3.3, Eq. (4), and Figure 5] The swing-up result is also measured through the teacher's own sensor loop. The reward in Eq. (4) is the FFT amplitude of the force sensor at the dominant frequency, while the teacher arm's helping motion is phase-aligned to the same estimated swing phase, with x_t = x0 + A_arm cos(theta_t). The paper reports no independent rope-angle, IMU-based tilt, or vision-based swing-height measurement, so the learned behavior could be an interaction with the arm rather than a true pendulum swing-up. Please add an independent swing-angle or swing-height trace, or otherwise validate quantitatively that the force amplitude tracks the actual swing amplitude.
- [Appendix C.5] Appendix C.5, labeled 'Real-world Learning Details', is an empty heading with no content. This is exactly the section that would document the real-world protocol for both tasks, including reset procedures, batch timing, evaluation protocol, and the measurements behind the 20-minute and 15-minute claims. Its absence makes the headline experiments difficult to reproduce and prevents the reader from assessing the proxy-reward concerns above. The section should be filled in or its content should be integrated into the main experimental details.
- [Figure 4 and Section 4.1] The walking ablations are evaluated with the same proxy reward that is being optimized. The y-axis 'linear velocity tracking rewards' is computed from the Eq. (9) treadmill speed, so the training curves conflate genuine policy improvement with exploitation of the force and pitch feedback loop. The 'better data efficiency' conclusion should be re-stated in terms of the measured walking outcome once an independent speed measurement is added; as written, the conclusion is about proxy reward rather than gait speed.
minor comments (5)
- [Appendix D] The heading 'Rapad Motor Adaptation' should be 'Rapid Motor Adaptation'.
- [Appendix C.2] The text refers to 'appliance control'; from the context this appears to be a typo for 'admittance control'.
- [Section 4.2] The sentence 'we apply an larger entropy coefficient' should be 'we apply a larger entropy coefficient'.
- [Section 3.3] The expression 'theta_0 = 30 degrees' uses an unusual degree symbol; use the standard notation for consistency.
- [References] Reference [10], a market research report, is an unusual source for the claim about industrial arm payload capacity; a technical datasheet or manufacturer specification would be more appropriate.
Circularity Check
No significant circularity: the walking and swing-up claims are real-world RL measurements with ablations; the treadmill/force proxy is a measurement-validity caveat, not a derivation tautology.
full rationale
Walking the derivation chain, the paper does not derive a headline result from equations that already contain that result. The three-stage latent optimization is a standard PPO/FiLM pipeline: a dynamics latent is optimized in simulation and then fine-tuned in the real world, and the claimed gains are empirical learning curves, not algebraic consequences of the reward definition. The walking reward uses the treadmill speed as a proxy for robot velocity (Eq. 3, with v approximated by the treadmill speed), and Eq. 9 defines that treadmill speed as a PD output depending on tether force and torso pitch; in principle, a policy could inflate its own reward by pulling on the tether. Similarly, the swing-up reward (Eq. 4) is the FFT amplitude of the force sensor that the teacher arm can directly excite during helping. These are serious measurement-channel/proxy concerns that belong to correctness and reproducibility review, and the paper's Section 6 explicitly acknowledges the lack of ground-reaction force sensing and proposes a force plate as future work. However, they are not circular in the required sense: no fitted parameter is renamed as a prediction, no result is mathematically forced by self-citation, and no uniqueness theorem is imported from the authors' prior work. Self-citations ([9], [35], [55]) provide the open-source hardware, encoder architecture, and three-stage training template; they are implementation support rather than load-bearing evidence for the central claim. The empty Appendix C.5 removes protocol detail and weakens verification, but missing detail is not circularity. The balance of evidence—ablations in Figure 4, FiLM ablations in Appendix B, and RMA comparison in Appendix D—gives the central claims independent empirical content. Score 2 reflects minor self-citations and the proxy-reward caveat, not a circular derivation chain.
Assumptions & free parameters
free parameters (7)
- walking reward shaping sigma =
100
- swing-up reward shaping alpha and target amplitude A_target =
alpha = 0.005, A_target approx mg*theta0 with m = 3.5 kg, theta0 = 30 deg
- arm guidance/perturbation amplitude A_arm =
0.05 m
- arm Z schedule =
linear decrease of 0.02 m over 5e4 env steps
- treadmill PD gains =
k1^p = 0.2, k2^p = -5, v_base = 0.1 m/s, speed cap 0.24 m/s
- FiLM layer learning rate =
5e-5
- dynamics latent z (universal and real-world) =
1024-dim vector; values not reported
assumptions (4)
- domain assumption Small-angle pendulum approximation for the swing-up target amplitude A_target approx m*g*theta0
- domain assumption Treadmill speed proxies the humanoid's true walking velocity
- domain assumption Elastic-rope tether permits free planar motion with smooth force transmission
- domain assumption The domain-randomized simulator is a valid training ground for the walking policy
invented entities (1)
-
dynamics latent z (FiLM-conditioned environment code)
independent evidence
Cite this review
Pith. "Pith review of Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids." pith.science (2026). https://pith.science/paper/RX6MALDH
@misc{pith2026250812252,
author = {Pith},
title = {Pith review of: Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids},
year = {2026},
howpublished = {\url{https://pith.science/paper/RX6MALDH}},
note = {Machine review of arXiv:2508.12252}
}
read the original abstract
Simulation-based reinforcement learning (RL) has significantly advanced humanoid locomotion tasks, yet direct real-world RL from scratch or adapting from pretrained policies remains rare, limiting the full potential of humanoid robots. Real-world learning, despite being crucial for overcoming the sim-to-real gap, faces substantial challenges related to safety, reward design, and learning efficiency. To address these limitations, we propose Robot-Trains-Robot (RTR), a novel framework where a robotic arm teacher actively supports and guides a humanoid robot student. The RTR system provides protection, learning schedule, reward, perturbation, failure detection, and automatic resets. It enables efficient long-term real-world humanoid training with minimal human intervention. Furthermore, we propose a novel RL pipeline that facilitates and stabilizes sim-to-real transfer by optimizing a single dynamics-encoded latent variable in the real world. We validate our method through two challenging real-world humanoid tasks: fine-tuning a walking policy for precise speed tracking and learning a humanoid swing-up task from scratch, illustrating the promising capabilities of real-world humanoid learning realized by RTR-style systems. See https://robot-trains-robot.github.io/ for more info.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
T. He, J. Gao, W. Xiao, Y . Zhang, Z. Wang, J. Wang, Z. Luo, G. He, N. Sobanbab, C. Pan, Z. Yi, G. Qu, K. Kitani, J. Hodgins, L. J. Fan, Y . Zhu, C. Liu, and G. Shi. ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills, Feb. 2025
work page 2025
-
[2]
T. He, W. Xiao, T. Lin, Z. Luo, Z. Xu, Z. Jiang, J. Kautz, C. Liu, G. Shi, X. Wang, L. Fan, and Y . Zhu. HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots, Mar. 2025
work page 2025
-
[3]
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath. Real-world humanoid locomotion with reinforcement learning. Science Robotics, 9(89):eadi9579, Apr. 2024. doi: 10.1126/scirobotics.adi9579
-
[4]
I. Radosavovic, S. Kamat, T. Darrell, and J. Malik. Learning Humanoid Locomotion over Challenging Terrain, Oct. 2024
work page 2024
- [5]
-
[6]
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 23–30, Sept. 2017. doi: 10.1109/IROS.2017.8202133
arXiv 2017
- [7]
-
[8]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms, 2017. URL https://arxiv.org/abs/1707.06347
arXiv 2017
Show all 58 references
-
[9]
H. Shi, W. Wang, S. Song, and C. K. Liu. ToddlerBot: Open-Source ML-Compatible Hu- manoid Platform for Loco-Manipulation, Feb. 2025
2025
-
[10]
Industrial robotic arm market size, share, growth trends, regional share, competitive intelligence, forecast report 2025–2037, March 2025
Research Nester. Industrial robotic arm market size, share, growth trends, regional share, competitive intelligence, forecast report 2025–2037, March 2025. URL https://www. researchnester.com/reports/industrial-robotic-arm-market/6763 . Accessed: 2025-04-28
2025
-
[11]
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel. Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 3803–3810, May 2018. doi: 10.1109/ICRA.2018.8460528
2018
-
[12]
Z. Yuan, T. Wei, S. Cheng, G. Zhang, Y . Chen, and H. Xu. Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning, Oct. 2024
2024
-
[13]
H. Ha, Y . Gao, Z. Fu, J. Tan, and S. Song. UMI-on-Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers. In 8th Annual Conference on Robot Learning, Sept. 2024
2024
-
[14]
Rudin, D
N. Rudin, D. Hoeller, P. Reist, and M. Hutter. Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning. In Proceedings of the 5th Conference on Robot Learn- ing, pages 91–100. PMLR, Jan. 2022
2022
-
[15]
Cheng, K
X. Cheng, K. Shi, A. Agarwal, and D. Pathak. Extreme Parkour with Legged Robots. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 11443–11450, May
2024
-
[16]
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y . Bai, D. Hafner, S. Bohez, and V . Vanhoucke. Sim-to-Real: Learning Agile Locomotion For Quadruped Robots. In Robotics: Science and Systems XIV, volume 14, June 2018. ISBN 978-0-9923747-4-7. 10
2018
-
[17]
F. Shi, Y . Kojio, T. Makabe, T. Anzai, K. Kojima, K. Okada, and M. Inaba. Reference- free learning bipedal motor skills via assistive force curricula. In A. Billard, T. Asfour, and O. Khatib, editors, Robotics Research, pages 304–320, Cham, 2023. Springer Nature Switzer- land...
2023
-
[18]
Akkaya, M
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang. Solving Rubik’s Cube with a Robot Hand, Oct. 2019
2019
-
[19]
T. Chen, M. Tippur, S. Wu, V . Kumar, E. Adelson, and P. Agrawal. Visual dexterity: In-hand reorientation of novel and complex object shapes. Science Robotics , 8(84):eadc9244, Nov
-
[20]
Y . Chen, C. Wang, L. Fei-Fei, and K. Liu. Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation. In 7th Annual Conference on Robot Learning , Aug. 2023
2023
-
[21]
Y . Chen, C. Wang, Y . Yang, and K. Liu. Object-Centric Dexterous Manipulation from Human Motion Data. In 8th Annual Conference on Robot Learning , Sept. 2024
2024
-
[22]
T. Lin, K. Sachdev, L. Fan, J. Malik, and Y . Zhu. Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids, Feb. 2025
2025
-
[23]
Z. Fu, Q. Zhao, Q. Wu, G. Wetzstein, and C. Finn. HumanPlus: Humanoid Shadowing and Imitation from Humans. In 8th Annual Conference on Robot Learning , Sept. 2024
2024
-
[24]
Gu, Y .-J
X. Gu, Y .-J. Wang, X. Zhu, C. Shi, Y . Guo, Y . Liu, and J. Chen. Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning. In Robotics: Science and Systems XX . Robotics: Science and Systems Foundation, July 2024. ISBN 9798990284807. ...
2024 doi
-
[25]
Haarnoja, B
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, J. Humplik, M. Wulfmeier, S. Tun- yasuvunakool, N. Y . Siegel, R. Hafner, M. Bloesch, K. Hartikainen, A. Byravan, L. Hasen- clever, Y . Tassa, F. Sadeghi, N. Batchelor, F. Casarini, S. Saliceti, C. Game, N. Sreendra, K...
2024 doi
-
[26]
Chebotar, A
Y . Chebotar, A. Handa, V . Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox. Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience. In 2019 International Conference on Robotics and Automation (ICRA) , pages 8973–8979, May
2019
-
[27]
A. Z. Ren, H. Dai, B. Burchfiel, and A. Majumdar. AdaptSim: Task-Driven Simulation Adap- tation for Sim-to-Real Transfer. In Proceedings of The 7th Conference on Robot Learning , pages 3434–3452. PMLR, Dec. 2023
2023
-
[28]
Huang, X
P. Huang, X. Zhang, Z. Cao, S. Liu, M. Xu, W. Ding, J. Francis, B. Chen, and D. Zhao. What Went Wrong? Closing the Sim-to-Real Gap via Differentiable Causal Discovery. In Proceedings of The 7th Conference on Robot Learning , pages 734–760. PMLR, Dec. 2023
2023
-
[29]
Schoettler, A
G. Schoettler, A. Nair, J. A. Ojea, S. Levine, and E. Solowjow. Meta-Reinforcement Learning for Robotic Industrial Insertion Tasks. In 2020 IEEE/RSJ International Conference on Intel- ligent Robots and Systems (IROS) , pages 9728–9735, Oct. 2020. doi: 10.1109/IROS45743. 2020.9340848
2020
-
[30]
Kumar, Z
A. Kumar, Z. Fu, D. Pathak, and J. Malik. RMA: Rapid Motor Adaptation for Legged Robots, July 2021. 11
2021
-
[31]
H. Qi, A. Kumar, R. Calandra, Y . Ma, and J. Malik. In-Hand Object Rotation via Rapid Motor Adaptation. In 6th Annual Conference on Robot Learning , Aug. 2022
2022
-
[32]
Kumar, Z
A. Kumar, Z. Li, J. Zeng, D. Pathak, K. Sreenath, and J. Malik. Adapting Rapid Motor Adap- tation for Bipedal Robots. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1161–1168, Oct. 2022. doi: 10.1109/IROS47612.2022.9981091
2022
-
[33]
Y . Sun, W. L. Ubellacker, W.-L. Ma, X. Zhang, C. Wang, N. V . Csomay-Shanklin, M. Tomizuka, K. Sreenath, and A. D. Ames. Online Learning of Unknown Dynamics for Model-Based Controllers in Legged Locomotion. IEEE Robotics and Automation Letters , 6 (4):8442–8449, Oct. 2021. IS...
2021
-
[34]
Zhang, L
Y . Zhang, L. Ke, A. Deshpande, A. Gupta, and S. Srinivasa. Cherry-Picking with Reinforce- ment Learning. In Robotics: Science and Systems XIX , volume 19, July 2023. ISBN 978-0- 9923747-9-2
2023
-
[35]
K. Lei, Z. He, C. Lu, K. Hu, Y . Gao, and H. Xu. Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization. InThe Twelfth International Conference on Learning Representations, Oct. 2023
2023
-
[36]
Xiong, R
H. Xiong, R. Mendonca, K. Shaw, and D. Pathak. Adaptive Mobile Manipulation for Articu- lated Objects In the Open World, Jan. 2024
2024
-
[37]
Jiang, C
Y . Jiang, C. Wang, R. Zhang, J. Wu, and L. Fei-Fei. TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction. In 8th Annual Conference on Robot Learning , Sept. 2024
2024
-
[38]
J. Luo, C. Xu, J. Wu, and S. Levine. Precise and Dexterous Robotic Manipulation via Human- in-the-Loop Reinforcement Learning, Mar. 2025
2025
-
[39]
Gupta, R
A. Gupta, R. Mendonca, Y . Liu, P. Abbeel, and S. Levine. Meta-reinforcement learning of structured exploration strategies, 2018. URL https://arxiv.org/abs/1802.07245
2018 arXiv
-
[40]
Rakelly, A
K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine. Efficient off-policy meta- reinforcement learning via probabilistic context variables, 2019. URL https://arxiv.org/ abs/1903.08254
2019 arXiv
-
[41]
W. Yu, J. Tan, Y . Bai, E. Coumans, and S. Ha. Learning fast adaptation with meta strategy optimization, 2020. URL https://arxiv.org/abs/1909.12995
2020 arXiv
-
[42]
X. B. Peng, E. Coumans, T. Zhang, T.-W. Lee, J. Tan, and S. Levine. Learning agile robotic locomotion skills by imitating animals, 2020. URL https://arxiv.org/abs/2004.00784
2020 arXiv
-
[43]
W. Yu, V . C. Kumar, G. Turk, and C. K. Liu. Sim-to-real transfer for biped locomotion, 2019. URL https://arxiv.org/abs/1903.01390
2019 arXiv
-
[44]
Huang, Z
S. Huang, Z. Zhang, T. Liang, Y . Xu, Z. Kou, C. Lu, G. Xu, Z. Xue, and H. Xu. MEN- TOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning, Oct. 2024
2024
-
[45]
K. Xu, Z. Hu, R. Doshi, A. Rovinsky, V . Kumar, A. Gupta, and S. Levine. Dexterous Ma- nipulation from Images: Autonomous Real-World RL via Substep Guidance. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 5938–5945, May 2023. doi: 10.1109/ICRA4...
2023
-
[46]
Mendonca, E
R. Mendonca, E. Panov, B. Bucher, J. Wang, and D. Pathak. Continuously Improving Mobile Manipulation with Autonomous Real-World RL. In Proceedings of The 8th Conference on Robot Learning, pages 5204–5219. PMLR, Jan. 2025. 12
2025
-
[47]
S. Ha, P. Xu, Z. Tan, S. Levine, and J. Tan. Learning to Walk in the Real World with Minimal Human Effort. In Proceedings of the 2020 Conference on Robot Learning , pages 1110–1120. PMLR, Oct. 2021
2020
-
[48]
Smith, I
L. Smith, I. Kostrikov, and S. Levine. A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning, Aug. 2022
2022
-
[49]
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg. DayDreamer: World Models for Physical Robot Learning. In Proceedings of The 6th Conference on Robot Learning , pages 2226–2240. PMLR, Mar. 2023
2023
-
[50]
Smith, Y
L. Smith, Y . Cao, and S. Levine. Grow your limits: Continuous improvement with real-world rl for robotic locomotion, 2023. URL https://arxiv.org/abs/2310.17634
2023 arXiv
-
[51]
Bloesch, J
M. Bloesch, J. Humplik, V . Patraucean, R. Hafner, T. Haarnoja, A. Byravan, N. Y . Siegel, S. Tunyasuvunakool, F. Casarini, N. Batchelor, F. Romano, S. Saliceti, M. Riedmiller, S. M. A. Eslami, and N. Heess. Towards real robot learning in the wild: A case study in bipedal loco...
2022
-
[52]
ROBOTIS OP3 e-Manual
ROBOTIS. ROBOTIS OP3 e-Manual. https://emanual.robotis.com/docs/en/ platform/op3/introduction/, 2024. Accessed: 2025-07-29
2024
-
[53]
Maples and J
J. Maples and J. Becker. Experiments in force control of robotic manipulators. In 1986 IEEE International Conference on Robotics and Automation Proceedings, volume 3, pages 695–702, Apr. 1986. doi: 10.1109/ROBOT.1986.1087590
1986
-
[54]
Y . Hou, Z. Liu, C. Chi, E. Cousineau, N. Kuppuswamy, S. Feng, B. Burchfiel, and S. Song. Adaptive Compliance Policy: Learning Approximate Compliance for Diffusion Guided Con- trol, Mar. 2025
2025
-
[55]
Ground Truth
Y . Liang, T. Xu, K. Hu, G. Jiang, F. Huang, and H. Xu. Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , Nov. 2024. 13 A State and Action Representation A...
2024
-
[2019]
doi: 10.1109/ICRA.2019.8793789
2019
-
[2023]
doi: 10.1126/scirobotics.adc9244
-
[2024]
doi: 10.1109/ICRA57147.2024.10610200
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.