Pith. sign in

REVIEW 8 cited by

From Imitation to Refinement -- Residual RL for Precise Assembly

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.16677 v4 pith:EXYZIGTH submitted 2024-07-23 cs.RO cs.LG

classification cs.ROcs.LG
keywords closed-loopactiondataperformanceresidualchunkscriticaldistribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in Behavior Cloning (BC) have made it easy to teach robots new tasks. However, we find that the ease of teaching comes at the cost of unreliable performance that saturates with increasing data for tasks requiring precision. The performance saturation can be attributed to two critical factors: (a) distribution shift resulting from the use of offline data and (b) the lack of closed-loop corrective control caused by action chucking (predicting a set of future actions executed open-loop) critical for BC performance. Our key insight is that by predicting action chunks, BC policies function more like trajectory "planners" than closed-loop controllers necessary for reliable execution. To address these challenges, we devise a simple yet effective method, ResiP (Residual for Precise Manipulation), that overcomes the reliability problem while retaining BC's ease of teaching and long-horizon capabilities. ResiP augments a frozen, chunked BC model with a fully closed-loop residual policy trained with reinforcement learning (RL) that addresses distribution shifts and introduces closed-loop corrections over open-loop execution of action chunks predicted by the BC trajectory planner. Videos, code, and data: https://residual-assembly.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0 of 10

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  2. The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    For high-precision manipulation, required demonstration count grows as log(N) ∝ 1/(P−c), where the fitted c varies with sensors, expert, and task complexity.

  3. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  4. Fabrica: Dual-Arm Assembly of General Multi-Part Objects via Integrated Planning and Learning

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A dual-arm robotic system combining hierarchical planning with equivariant residual RL policies demonstrates multi-part assembly of five-to-nine-part objects, with strong step-level but weaker end-to-end real-world success.

  5. Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    Approximation error of constant-depth parallelizable sequence models falls exponentially with depth, via a Lie-algebraic tower of expressivity extensions.

  6. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

  7. GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective

    cs.AI 2025-07 unverdicted novelty 3.0 of 10

    A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.

  8. A Survey on Imitation Learning for Contact-Rich Tasks in Robotics

    cs.RO 2025-06

Pith tools