REVIEW 6 cited by
Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider the problem of learning useful robotic skills from previously collected offline data without access to manually specified rewards or additional online exploration, a setting that is becoming increasingly important for scaling robot learning by reusing past robotic data. In particular, we propose the objective of learning a functional understanding of the environment by learning to reach any goal state in a given dataset. We employ goal-conditioned Q-learning with hindsight relabeling and develop several techniques that enable training in a particularly challenging offline setting. We find that our method can operate on high-dimensional camera images and learn a variety of skills on real robots that generalize to previously unseen scenes and objects. We also show that our method can learn to reach long-horizon goals across multiple episodes through goal chaining, and learn rich representations that can help with downstream tasks through pre-training or auxiliary objectives. The videos of our experiments can be found at https://actionable-models.github.io
Forward citations
Cited by 6 Pith papers
-
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
CoGHP autoregressively plans a chain of latent subgoals inside a single MLP-Mixer policy and beats prior hierarchical offline RL baselines on most OGBench tasks.
-
Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
VfO trains a state-value function on action-free expert demonstrations mixed with lower-quality background data, then uses advantage-weighted regression on the background data to improve the agent, approaching oracle ...
-
Reachability Weighted Offline Goal-conditioned Resampling
RWS trains a positive-unlabeled reachability classifier on goal-conditioned Q-values and uses it to re-weight goal sampling, improving offline goal-conditioned RL performance on robotic manipulation benchmarks.
-
Prompting Decision Transformers for Zero-Shot Reach-Avoid Policies
RADT learns reach-avoid policies from random offline trajectories and generalizes to unseen avoid-region sizes and counts via prompt conditioning.
-
Relative Value Learning
A critic that learns antisymmetric value differences ∆(s_i,s_j)=V(s_i)−V(s_j) has a provably contracting Bellman operator and an unbiased advantage estimator, and PPO with this critic matches standard PPO on Atari.
-
Mollified Value Learning
Mollified Value Learning regularizes offline goal-conditioned value estimates with a Feynman-Kac expectation version of the viscous HJB equation instead of a pointwise Eikonal constraint.
Discussion (0). Continue with ORCID to comment.