Pith. sign in

REVIEW 8 cited by

Reward-Conditioned Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.13465 v1 pith:7V4U452U submitted 2019-12-31 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningmethodspoliciespolicyreinforcementrewardsupervisedmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning offers the promise of automating the acquisition of complex behavioral skills. However, compared to commonly used and well-understood supervised learning methods, reinforcement learning algorithms can be brittle, difficult to use and tune, and sensitive to seemingly innocuous implementation decisions. In contrast, imitation learning utilizes standard and well-understood supervised learning methods, but requires near-optimal expert data. Can we learn effective policies via supervised learning without demonstrations? The main idea that we explore in this work is that non-expert trajectories collected from sub-optimal policies can be viewed as optimal supervision, not for maximizing the reward, but for matching the reward of the given trajectory. By then conditioning the policy on the numerical value of the reward, we can obtain a policy that generalizes to larger returns. We show how such an approach can be derived as a principled method for policy search, discuss several variants, and compare the method experimentally to a variety of current reinforcement learning methods on standard benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffusion Guidance Is a Controllable Policy Improvement Operator

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Diffusion guidance with a tunable weight is a controllable policy improvement operator, improving offline and goal-conditioned policies beyond the data without retraining and often without a value function.

  2. A Provable Approach for End-to-End Safe Reinforcement Learning

    cs.LG 2025-05 conditional novelty 7.0 of 10

    PLS combines offline return-conditioned policy training with Gaussian-process safe optimization of target returns to provide high-probability safety throughout deployment.

  3. Freeform Preference Learning for Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.5 of 10

    Language-conditioned multi-axis human preferences yield denser rewards and steerable robot policies that outperform sparse and binary-preference baselines by 38 points on long-horizon manipulation.

  4. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  5. Generative Sequential Notification Optimization via Multi-Objective Decision Transformers

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A Decision Transformer with quantile-regression return prompts improved notification decisions at LinkedIn, boosting sessions by 0.72% over the deployed CQL baseline in a live A/B test.

  6. EBaReT: Expert-guided Bag Reward Transformer for Auto Bidding

    cs.LG 2025-07 conditional novelty 6.0 of 10

    EBaReT, a Decision Transformer variant with a PU-learning expert discriminator and bag-level reward redistribution, outperforms offline RL and generative baselines on seven auto-bidding benchmark periods.

  7. Behavioral Exploration: Learning to Explore via In-Context Adaptation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.

  8. How to Provably Improve Return Conditioned Supervised Learning?

    cs.LG 2025-06 conditional novelty 6.0 of 10

    R2CSL provably reaches the in-distribution optimal stitched policy by conditioning on the maximum return-to-go per state, improving on standard RCSL without dynamic programming.

Pith tools