Pith. sign in

REVIEW 10 cited by

Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02875 v2 pith:RFME6KYU submitted 2019-12-05 cs.AI cs.LG

classification cs.AIcs.LG
keywords udrlrewardscommandshumansimitateinputlearningrobot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We transform reinforcement learning (RL) into a form of supervised learning (SL) by turning traditional RL on its head, calling this Upside Down RL (UDRL). Standard RL predicts rewards, while UDRL instead uses rewards as task-defining inputs, together with representations of time horizons and other computable functions of historic and desired future data. UDRL learns to interpret these input observations as commands, mapping them to actions (or action probabilities) through SL on past (possibly accidental) experience. UDRL generalizes to achieve high rewards or other goals, through input commands such as: get lots of reward within at most so much time! A separate paper [63] on first experiments with UDRL shows that even a pilot version of UDRL can outperform traditional baseline algorithms on certain challenging RL problems. We also also conceptually simplify an approach [60] for teaching a robot to imitate humans. First videotape humans imitating the robot's current behaviors, then let the robot learn through SL to map the videos (as input commands) to these behaviors, then let it generalize and imitate videos of humans executing previously unknown behavior. This Imitate-Imitator concept may actually explain why biological evolution has resulted in parents who imitate the babbling of their babies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

    cs.AI 2026-01 accept novelty 7.0 of 10

    Success conditioning exactly maximizes linearized policy improvement under a chi-squared divergence trust region whose radius is action-influence, with relative improvement, policy change, and action-influence exactly...

  2. Single-pass Adaptive Image Tokenization for Minimum Program Search

    cs.CV 2025-07 conditional novelty 7.0 of 10

    KARL conditions a tokenizer on a target reconstruction loss and learns halting probabilities that produce an adaptive token count in a single forward pass.

  3. BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL

    cs.LG 2025-05 conditional novelty 7.0 of 10

    A learned Transformer-based acquisition policy for multi-objective Bayesian optimization, trained on synthetic Gaussian processes, achieves best or near-best hypervolume on most tested synthetic and 3D Gaussian Splatt...

  4. A Provable Approach for End-to-End Safe Reinforcement Learning

    cs.LG 2025-05 conditional novelty 7.0 of 10

    PLS combines offline return-conditioned policy training with Gaussian-process safe optimization of target returns to provide high-probability safety throughout deployment.

  5. Freeform Preference Learning for Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.5 of 10

    Language-conditioned multi-axis human preferences yield denser rewards and steerable robot policies that outperform sparse and binary-preference baselines by 38 points on long-horizon manipulation.

  6. Behavioral Exploration: Learning to Explore via In-Context Adaptation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.

  7. How to Provably Improve Return Conditioned Supervised Learning?

    cs.LG 2025-06 conditional novelty 6.0 of 10

    R2CSL provably reaches the in-distribution optimal stitched policy by conditioning on the maximum return-to-go per state, improving on standard RCSL without dynamic programming.

  8. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.

  9. GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A goal-agnostic curiosity reward is claimed to make active geo-localization agents generalize better to unseen targets and environments than distance-based rewards.

  10. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools