Pith. sign in

REVIEW 6 cited by

Residual Reinforcement Learning from Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.08050 v1 pith:OD7HCS2E submitted 2021-06-15 cs.LG

classification cs.LG
keywords demonstrationsresidualcontrollerlearningtasksinputsreinforcementreward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Residual reinforcement learning (RL) has been proposed as a way to solve challenging robotic tasks by adapting control actions from a conventional feedback controller to maximize a reward signal. We extend the residual formulation to learn from visual inputs and sparse rewards using demonstrations. Learning from images, proprioceptive inputs and a sparse task-completion reward relaxes the requirement of accessing full state features, such as object and target positions. In addition, replacing the base controller with a policy learned from demonstrations removes the dependency on a hand-engineered controller in favour of a dataset of demonstrations, which can be provided by non-experts. Our experimental evaluation on simulated manipulation tasks on a 6-DoF UR5 arm and a 28-DoF dexterous hand demonstrates that residual RL from demonstrations is able to generalize to unseen environment conditions more flexibly than either behavioral cloning or RL fine-tuning, and is capable of solving high-dimensional, sparse-reward tasks out of reach for RL from scratch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video

    cs.RO 2026-08 conditional novelty 7.0 of 10

    C2Dex converts monocular human videos into executable dexterous robot manipulation trajectories by using stable object-side contacts as a shared representation for reconstruction and retargeting, achieving 57.78% and ...

  2. Touch begins where vision ends: Generalizable policies for contact-rich manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A localize-then-execute policy that combines vision-language reaching, semantic background augmentation, and residual reinforcement learning with tactile sensing reaches about 90% success on millimeter-precision manip...

  3. Learning Residual Kinematic Corrections for Continuous Neural Decoding via Reinforcement Learning

    cs.AI 2026-07 conditional novelty 5.5 of 10

    Offline residual SAC correction on CNN–LSTM kinematic outputs improves continuous 3D EEG motor-imagery decoding by ~21–42% in correlation and RMSE across 2D and VR feedback without feeding EEG to the RL agent.

  4. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

  5. ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills

    cs.RO 2025-02 conditional novelty 5.0 of 10

    ASAP trains a residual action model on real-world rollouts and fine-tunes simulation policies through it, reducing humanoid whole-body motion tracking error in sim-to-real transfer.

  6. Learning-based Autonomous Oversteer Control and Collision Avoidance

    cs.RO 2025-05 conditional novelty 4.0 of 10

    QC-SAC reaches 81.8% success on a new simulated oversteer-plus-obstacle-avoidance benchmark, beating BC, SAC, and BC-SAC baselines.

Pith tools