REVIEW 10 cited by
Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We transform reinforcement learning (RL) into a form of supervised learning (SL) by turning traditional RL on its head, calling this Upside Down RL (UDRL). Standard RL predicts rewards, while UDRL instead uses rewards as task-defining inputs, together with representations of time horizons and other computable functions of historic and desired future data. UDRL learns to interpret these input observations as commands, mapping them to actions (or action probabilities) through SL on past (possibly accidental) experience. UDRL generalizes to achieve high rewards or other goals, through input commands such as: get lots of reward within at most so much time! A separate paper [63] on first experiments with UDRL shows that even a pilot version of UDRL can outperform traditional baseline algorithms on certain challenging RL problems. We also also conceptually simplify an approach [60] for teaching a robot to imitate humans. First videotape humans imitating the robot's current behaviors, then let the robot learn through SL to map the videos (as input commands) to these behaviors, then let it generalize and imitate videos of humans executing previously unknown behavior. This Imitate-Imitator concept may actually explain why biological evolution has resulted in parents who imitate the babbling of their babies.
Forward citations
Cited by 10 Pith papers
-
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success
Success conditioning exactly maximizes linearized policy improvement under a chi-squared divergence trust region whose radius is action-influence, with relative improvement, policy change, and action-influence exactly...
-
Single-pass Adaptive Image Tokenization for Minimum Program Search
KARL conditions a tokenizer on a target reconstruction loss and learns halting probabilities that produce an adaptive token count in a single forward pass.
-
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
A learned Transformer-based acquisition policy for multi-objective Bayesian optimization, trained on synthetic Gaussian processes, achieves best or near-best hypervolume on most tested synthetic and 3D Gaussian Splatt...
-
A Provable Approach for End-to-End Safe Reinforcement Learning
PLS combines offline return-conditioned policy training with Gaussian-process safe optimization of target returns to provide high-probability safety throughout deployment.
-
Freeform Preference Learning for Robotic Manipulation
Language-conditioned multi-axis human preferences yield denser rewards and steerable robot policies that outperform sparse and binary-preference baselines by 38 points on long-horizon manipulation.
-
Behavioral Exploration: Learning to Explore via In-Context Adaptation
A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.
-
How to Provably Improve Return Conditioned Supervised Learning?
R2CSL provably reaches the in-distribution optimal stitched policy by conditioning on the maximum return-to-go per state, improving on standard RCSL without dynamic programming.
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.
-
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
A goal-agnostic curiosity reward is claimed to make active geo-localization agents generalize better to unseen targets and environments than distance-based rewards.
-
Reinforcement Learning: From Algorithms To Foundation Models
A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.
Discussion (0). Continue with ORCID to comment.