Pith. sign in

REVIEW 6 cited by

Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.07749 v3 pith:CKJ2QTHB submitted 2021-04-15 cs.RO cs.LG

classification cs.ROcs.LG
keywords learninglearnofflineroboticskillsdatagoalmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider the problem of learning useful robotic skills from previously collected offline data without access to manually specified rewards or additional online exploration, a setting that is becoming increasingly important for scaling robot learning by reusing past robotic data. In particular, we propose the objective of learning a functional understanding of the environment by learning to reach any goal state in a given dataset. We employ goal-conditioned Q-learning with hindsight relabeling and develop several techniques that enable training in a particularly challenging offline setting. We find that our method can operate on high-dimensional camera images and learn a variety of skills on real robots that generalize to previously unseen scenes and objects. We also show that our method can learn to reach long-horizon goals across multiple episodes through goal chaining, and learn rich representations that can help with downstream tasks through pre-training or auxiliary objectives. The videos of our experiments can be found at https://actionable-models.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL

    cs.LG 2026-02 conditional novelty 6.0 of 10

    CoGHP autoregressively plans a chain of latent subgoals inside a single MLP-Mixer policy and beats prior hierarchical offline RL baselines on most OGBench tasks.

  2. Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement

    cs.LG 2025-07 conditional novelty 6.0 of 10

    VfO trains a state-value function on action-free expert demonstrations mixed with lower-quality background data, then uses advantage-weighted regression on the background data to improve the agent, approaching oracle ...

  3. Reachability Weighted Offline Goal-conditioned Resampling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    RWS trains a positive-unlabeled reachability classifier on goal-conditioned Q-values and uses it to re-weight goal sampling, improving offline goal-conditioned RL performance on robotic manipulation benchmarks.

  4. Prompting Decision Transformers for Zero-Shot Reach-Avoid Policies

    cs.LG 2025-05 conditional novelty 6.0 of 10

    RADT learns reach-avoid policies from random offline trajectories and generalizes to unseen avoid-region sizes and counts via prompt conditioning.

  5. Relative Value Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A critic that learns antisymmetric value differences ∆(s_i,s_j)=V(s_i)−V(s_j) has a provably contracting Bellman operator and an unbiased advantage estimator, and PPO with this critic matches standard PPO on Atari.

  6. Mollified Value Learning

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Mollified Value Learning regularizes offline goal-conditioned value estimates with a Feynman-Kac expectation version of the viscous HJB equation instead of a pointwise Eikonal constraint.

Pith tools