Pith. sign in

REVIEW 2 cited by

Scaling data-driven robotics with reward sketching and batch reinforcement learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.12200 v3 pith:HVRD2KLN submitted 2019-09-26 cs.RO cs.LG

classification cs.ROcs.LG
keywords tasksrewardexperiencerobotbatchdata-drivendatasetdifferent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a framework for data-driven robotics that makes use of a large dataset of recorded robot experience and scales to several tasks using learned reward functions. We show how to apply this framework to accomplish three different object manipulation tasks on a real robot platform. Given demonstrations of a task together with task-agnostic recorded experience, we use a special form of human annotation as supervision to learn a reward function, which enables us to deal with real-world tasks where the reward signal cannot be acquired directly. Learned rewards are used in combination with a large dataset of experience from different tasks to learn a robot policy offline using batch RL. We show that using our approach it is possible to train agents to perform a variety of challenging manipulation tasks including stacking rigid objects and handling cloth.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Guiding Data Collection via Factored Scaling Curves

    cs.RO 2025-05 conditional novelty 7.0 of 10

    Factored scaling curves that rank environmental factors by predicted marginal success gain allocate a fixed robot data budget more effectively than equal, greedy, or robust-mixture baselines in simulation and real-wor...

  2. SenDaL: An Effective and Efficient Calibration Framework of Low-Cost Sensors for Daily Life

    cs.LG 2025-02 conditional novelty 4.0 of 10

    SenDaL trains a router to switch between a linear and a deep calibration model, achieving deep-model accuracy at near-linear-model speed on low-cost fine-dust sensors.

Pith tools