Pith. sign in

REVIEW 8 cited by

The Ingredients of Real-World Robotic Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.12570 v1 pith:6C6VF55M submitted 2020-04-27 cs.LG cs.ROstat.ML

classification cs.LGcs.ROstat.ML
keywords learningsystemwithoutchallengesroboticdemonstratedexteroushuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The success of reinforcement learning for real world robotics has been, in many cases limited to instrumented laboratory scenarios, often requiring arduous human effort and oversight to enable continuous learning. In this work, we discuss the elements that are needed for a robotic learning system that can continually and autonomously improve with data collected in the real world. We propose a particular instantiation of such a system, using dexterous manipulation as our case study. Subsequently, we investigate a number of challenges that come up when learning without instrumentation. In such settings, learning must be feasible without manually designed resets, using only on-board perception, and without hand-engineered reward functions. We propose simple and scalable solutions to these challenges, and then demonstrate the efficacy of our proposed system on a set of dexterous robotic manipulation tasks, providing an in-depth analysis of the challenges associated with this learning paradigm. We demonstrate that our complete system can learn without any human intervention, acquiring a variety of vision-based skills with a real-world three-fingered hand. Results and videos can be found at https://sites.google.com/view/realworld-rl/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces

    cs.LG 2025-07 conditional novelty 7.0 of 10

    For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.

  2. From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning

    cs.RO 2026-03 accept novelty 6.0 of 10

    Residual off-policy RL with selective BC regularization and value-guided sampling contracts a pretrained generative robot policy around successful actions, reaching high success on hard long-horizon tasks from pixels ...

  3. Auto-exploration for online reinforcement learning

    cs.LG 2025-12 conditional novelty 6.0 of 10

    New parameter-free SPMD algorithms achieve the first algorithm-independent O(ε⁻²) sample complexity for online discounted RL under a mixing-optimal-policy assumption.

  4. SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SafeMimic enables a mobile robot to safely and autonomously adapt a single third-person human video into a successful multi-step manipulation strategy.

  5. CARoL: Context-aware Adaptation for Robot Learning

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CARoL measures task similarity by state-transition prediction errors and uses those similarities to weight prior policies, value functions, or actor-critic knowledge when adapting to a new task.

  6. Flow-based Domain Randomization for Learning and Sequencing Robotic Skills

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A normalizing-flow sampling distribution, trained with entropy-regularized reward maximization, improves domain coverage and sim-to-real transfer over Gaussian, beta, and interval-based learned domain randomization.

  7. SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Simulation-pretrained policies, with digital-twin demos for critic bootstrapping and action proposals, cut real-world RL training time while reaching near-perfect success on three manipulation tasks.

  8. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

Pith tools