Pith. sign in

REVIEW 11 cited by

MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08212 v2 pith:XWQLXNYS submitted 2021-04-16 cs.RO cs.LG

classification cs.ROcs.LG
keywords taskssystemlearningmt-optacquireexperienceframeworkreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

General-purpose robotic systems must master a large repertoire of diverse skills to be useful in a range of daily tasks. While reinforcement learning provides a powerful framework for acquiring individual behaviors, the time needed to acquire each skill makes the prospect of a generalist robot trained with RL daunting. In this paper, we study how a large-scale collective robotic learning system can acquire a repertoire of behaviors simultaneously, sharing exploration, experience, and representations across tasks. In this framework new tasks can be continuously instantiated from previously learned tasks improving overall performance and capabilities of the system. To instantiate this system, we develop a scalable and intuitive framework for specifying new tasks through user-provided examples of desired outcomes, devise a multi-robot collective learning system for data collection that simultaneously collects experience for multiple tasks, and develop a scalable and generalizable multi-task deep reinforcement learning method, which we call MT-Opt. We demonstrate how MT-Opt can learn a wide range of skills, including semantic picking (i.e., picking an object from a particular category), placing into various fixtures (e.g., placing a food item onto a plate), covering, aligning, and rearranging. We train and evaluate our system on a set of 12 real-world tasks with data collected from 7 robots, and demonstrate the performance of our system both in terms of its ability to generalize to structurally similar new tasks, and acquire distinct new tasks more quickly by leveraging past experience. We recommend viewing the videos at https://karolhausman.github.io/mt-opt/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ScanBot: A Benchmark for Precision Robotic Surface Scanning with Industrial Laser Profilers

    cs.RO 2025-05 conditional novelty 7.0 of 10

    Introduces the first instruction-conditioned benchmark for precision laser surface scanning and shows that current multimodal LLMs fail at parameter selection and region grounding.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  4. SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    SAFE-Pruner forecasts deep-layer visual-token saliency from historical attention maps and refreshes at subtask boundaries, enabling up to 1.89x faster VLA inference with minimal success-rate drop.

  5. TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

    cs.RO 2026-02 conditional novelty 6.0 of 10

    The log-probability a VLM assigns to 'True' for 'does this video prefix complete the task?' is used as a zero-shot dense progress reward that outperforms GVL on open-source models.

  6. EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Co-training a vision-language-action humanoid policy on aligned egocentric human demonstrations plus limited robot data improves real-world loco-manipulation success by 20% in-domain and 51% in environments the robot ...

  7. RAD: Retrieval High-quality Demonstrations to Enhance Decision-making

    cs.AI 2025-07 conditional novelty 6.0 of 10

    RAD retrieves high-return states from an offline dataset and uses condition-guided diffusion to plan toward them, reporting competitive D4RL MuJoCo scores.

  8. Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    GO-Skill learns a discrete library of goal-oriented skills from offline multi-task data and selects them with a hierarchical policy, improving average episode returns on MetaWorld MT30 and MT50.

  9. CARoL: Context-aware Adaptation for Robot Learning

    cs.RO 2025-06 conditional novelty 6.0 of 10

    CARoL measures task similarity by state-transition prediction errors and uses those similarities to weight prior policies, value functions, or actor-critic knowledge when adapting to a new task.

  10. Training RL Agents for Multi-Objective Network Defense Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Diverse, dynamically ordered training tasks make network-defense RL agents generalize to unseen attacks better than single-task training.

  11. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools