REVIEW 4 cited by
Reinforcement Learning for Robust Athletic Intelligence: Lessons from the 2nd 'AI Olympics with RealAIGym' Competition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In the field of robotics many different approaches ranging from classical planning over optimal control to reinforcement learning (RL) are developed and borrowed from other fields to achieve reliable control in diverse tasks. In order to get a clear understanding of their individual strengths and weaknesses and their applicability in real world robotic scenarios is it important to benchmark and compare their performances not only in a simulation but also on real hardware. The '2nd AI Olympics with RealAIGym' competition was held at the IROS 2024 conference to contribute to this cause and evaluate different controllers according to their ability to solve a dynamic control problem on an underactuated double pendulum system with chaotic dynamics. This paper describes the four different RL methods submitted by the participating teams, presents their performance in the swing-up task on a real double pendulum, measured against various criteria, and discusses their transferability from simulation to real hardware and their robustness to external disturbances.
Forward citations
Cited by 4 Pith papers
-
VIMPPI: Enhancing Model Predictive Path Integral Control with Variational Integration for Underactuated Systems
Using a variational integrator inside MPPI rollouts lets the controller plan 4-20 times further ahead, improving balance uptime on underactuated double pendulums.
-
Finetuning Deep Reinforcement Learning Policies with Evolutionary Strategies for Control of Underactuated Robots
Hybrid SAC+SNES training improves swing-up and competition scores for underactuated robots over RL-only baselines.
-
Real-Time Nonlinear MPC via Sequential Quadratic Programming with Structure-Exploiting ADMM and Interior-Point Methods for Underactuated Double-Pendulum Swing-Up
An SQP-based MPC controller with inverse-dynamics formulation swings up and stabilizes underactuated double pendulums on hardware, with 100% success in three of four scenarios but 70% on the disturbed Acrobot.
-
Average-Reward Maximum Entropy Reinforcement Learning for Global Policy in Double Pendulum Tasks
The authors' AR-EAPO controller achieves high simulated scores on swing-up tasks for acrobot and pendubot under increased disturbances by widening initial-state variance and shortening effective horizon during training.
Discussion (0). Continue with ORCID to comment.