Pith. sign in

REVIEW 2 cited by

Learning Active Task-Oriented Exploration Policies for Bridging the Sim-to-Real Gap

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.01952 v2 pith:FC7QROGD submitted 2020-06-02 cs.RO

classification cs.RO
keywords explorationpoliciesparametersdynamicslearningperformtasktask-oriented
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training robotic policies in simulation suffers from the sim-to-real gap, as simulated dynamics can be different from real-world dynamics. Past works tackled this problem through domain randomization and online system-identification. The former is sensitive to the manually-specified training distribution of dynamics parameters and can result in behaviors that are overly conservative. The latter requires learning policies that concurrently perform the task and generate useful trajectories for system identification. In this work, we propose and analyze a framework for learning exploration policies that explicitly perform task-oriented exploration actions to identify task-relevant system parameters. These parameters are then used by model-based trajectory optimization algorithms to perform the task in the real world. We instantiate the framework in simulation with the Linear Quadratic Regulator as well as in the real world with pouring and object dragging tasks. Experiments show that task-oriented exploration helps model-based policies adapt to systems with initially unknown parameters, and it leads to better task performance than task-agnostic exploration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning

    cs.RO 2025-05 conditional novelty 7.0 of 10

    SPI-Active identifies legged-robot physical parameters via massive parallel sampling and uses Fisher-information-optimal command sequences to collect informative real-world data, improving sim-to-real transfer on quad...

  2. Flow-based Domain Randomization for Learning and Sequencing Robotic Skills

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A normalizing-flow sampling distribution, trained with entropy-regularized reward maximization, improves domain coverage and sim-to-real transfer over Gaussian, beta, and interval-based learned domain randomization.

Pith tools