Pith. sign in

REVIEW 2 cited by

All You Need Is Supervised Learning: From Imitation Learning to Meta-RL With Upside Down RL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.11960 v1 pith:CRGJFHQ2 submitted 2022-02-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningudrldownsettingupsideagentimitationmeta-rl
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Upside down reinforcement learning (UDRL) flips the conventional use of the return in the objective function in RL upside down, by taking returns as input and predicting actions. UDRL is based purely on supervised learning, and bypasses some prominent issues in RL: bootstrapping, off-policy corrections, and discount factors. While previous work with UDRL demonstrated it in a traditional online RL setting, here we show that this single algorithm can also work in the imitation learning and offline RL settings, be extended to the goal-conditioned RL setting, and even the meta-RL setting. With a general agent architecture, a single UDRL agent can learn across all paradigms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Action Chunking with Transformers for Image-Based Spacecraft Guidance and Control

    cs.RO 2025-09 conditional novelty 6.0 of 10

    ACT imitation learning from 100 meta-RL demonstrations beats the meta-RL baseline on simulated ISS docking, using about 6,300 interactions instead of 40 million.

  2. Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A pessimism-based framework for zero-shot transfer RL builds conservative proxies from robust MDPs, yielding lower-bound performance guarantees and distributed algorithms that mitigate negative transfer.

Pith tools