Pith. sign in

REVIEW 11 cited by

Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21845 v3 pith:Q7K3K5ZM submitted 2024-10-29 cs.RO cs.AI

classification cs.ROcs.AI
keywords manipulationlearningpoliciesroboticapproachcomplexdexteroushuman-in-the-loop
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) holds great promise for enabling autonomous acquisition of complex robotic manipulation skills, but realizing this potential in real-world settings has been challenging. We present a human-in-the-loop vision-based RL system that demonstrates impressive performance on a diverse set of dexterous manipulation tasks, including dynamic manipulation, precision assembly, and dual-arm coordination. Our approach integrates demonstrations and human corrections, efficient RL algorithms, and other system-level design choices to learn policies that achieve near-perfect success rates and fast cycle times within just 1 to 2.5 hours of training. We show that our method significantly outperforms imitation learning baselines and prior RL approaches, with an average 2x improvement in success rate and 1.8x faster execution. Through extensive experiments and analysis, we provide insights into the effectiveness of our approach, demonstrating how it learns robust, adaptive policies for both reactive and predictive control strategies. Our results suggest that RL can indeed learn a wide range of complex vision-based manipulation policies directly in the real world within practical training times. We hope this work will inspire a new generation of learned robotic manipulation techniques, benefiting both industrial applications and research advancements. Videos and code are available at our project website https://hil-serl.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    For high-precision manipulation, required demonstration count grows as log(N) ∝ 1/(P−c), where the fitted c varies with sensors, expert, and task complexity.

  2. FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Human corrective actions can be inverted into noise-space targets that train a lightweight latent policy to steer frozen flow/diffusion robot models from a handful of interventions while preserving pretrained skills.

  3. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  4. RhinoVLA Technical Report

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    RhinoVLA uses a token-efficient Qwen3-VL backbone, continuous Action Expert, and unified cross-robot interface to match π0.5 performance while hitting 11.69 Hz on Huixi R1 edge SoC.

  5. $M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    M²-VLA shows that generalized VLMs can serve as direct backbones for robotic manipulation by selectively extracting task-critical features via Mixture of Layers and adding Meta Skill Modules for efficient trajectory learning.

  6. Prior Reinforce: Goal-Conditioned Dynamic Manipulation with Limited Trials

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Prior Reinforce adapts a few demonstration motions to new goals in dynamic manipulation by learning a diffusion motion prior and refining a low-dimensional condition via Bayesian optimization, reaching new goals in un...

  7. TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    TeViR uses a text-to-video diffusion model to generate future frames from a task description and rewards an RL agent for matching those frames, improving sample efficiency in robotic manipulation.

  8. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    A VLA policy with auxiliary success/progress heads and AWR+RECAP-style RL finished 1st in the LeHome 2026 simulation round and 2nd on the real robot.

  9. RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction

    cs.RO 2025-09 conditional novelty 5.0 of 10

    Robot policies trained on human interventions that rewind to a familiar state and then correct the mistake achieve higher long-horizon success and better data efficiency than imitation on full demonstrations alone.

  10. SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Simulation-pretrained policies, with digital-twin demos for critic bootstrapping and action proposals, cut real-world RL training time while reaching near-perfect success on three manipulation tasks.

  11. Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Breaking long manipulation tasks into atomic subtasks and collecting demonstrations from varied starting poses improves imitation learning success using fewer demonstration frames.

Pith tools