Pith. sign in

REVIEW 3 cited by

Grow Your Limits: Continuous Improvement with Real-World RL for Robotic Locomotion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17634 v1 pith:FKEO474P submitted 2023-10-26 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords trainingaprlcontinuedexplorationimprovementlimitslocomotionpolicy
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep reinforcement learning (RL) can enable robots to autonomously acquire complex behaviors, such as legged locomotion. However, RL in the real world is complicated by constraints on efficiency, safety, and overall training stability, which limits its practical applicability. We present APRL, a policy regularization framework that modulates the robot's exploration over the course of training, striking a balance between flexible improvement potential and focused, efficient exploration. APRL enables a quadrupedal robot to efficiently learn to walk entirely in the real world within minutes and continue to improve with more training where prior work saturates in performance. We demonstrate that continued training with APRL results in a policy that is substantially more capable of navigating challenging situations and is able to adapt to changes in dynamics with continued training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A force-sensing arm acts as teacher for a small humanoid, enabling 20-minute real-world walking speed adaptation and 15-minute swing-up learning from scratch.

  2. Towards Embodiment Scaling Laws in Robot Locomotion

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A policy trained on about one thousand simulated robot bodies generalizes progressively better to unseen bodies as the number of training bodies grows, and it transfers zero-shot to two real robots.

  3. Real Time Control of Tandem-Wing Experimental Platform Using Concerto Reinforcement Learning

    cs.LG 2025-02 reject novelty 4.0 of 10

    CRL2RT combines classical controllers with RL in a time-interleaved Cloud-Edge design, reporting over 2500 Hz online-update control on CPUs and tracking gains of 18.3% to 60.7% in simulation.

Pith tools