REVIEW 3 cited by
Synthetic Returns for Long-Term Credit Assignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Since the earliest days of reinforcement learning, the workhorse method for assigning credit to actions over time has been temporal-difference (TD) learning, which propagates credit backward timestep-by-timestep. This approach suffers when delays between actions and rewards are long and when intervening unrelated events contribute variance to long-term returns. We propose state-associative (SA) learning, where the agent learns associations between states and arbitrarily distant future rewards, then propagates credit directly between the two. In this work, we use SA-learning to model the contribution of past states to the current reward. With this model we can predict each state's contribution to the far future, a quantity we call "synthetic returns". TD-learning can then be applied to select actions that maximize these synthetic returns (SRs). We demonstrate the effectiveness of augmenting agents with SRs across a range of tasks on which TD-learning alone fails. We show that the learned SRs are interpretable: they spike for states that occur after critical actions are taken. Finally, we show that our IMPALA-based SR agent solves Atari Skiing -- a game with a lengthy reward delay that posed a major hurdle to deep-RL agents -- 25 times faster than the published state-of-the-art.
Forward citations
Cited by 3 Pith papers
-
Counterfactual Shapley Credit Assignment
Counterfactual Shapley values, computed by simulated 'what-if' action replacements, redistribute RL rewards without changing the optimal policy and improve credit assignment in stochastic, sparse, delayed-reward tasks.
-
Attention-Based Reward Shaping for Sparse and Delayed Rewards
ARES uses attention weights from a return-predicting transformer to generate dense shaped rewards from fully delayed reward episodes, improving RL training in many test environments.
-
Zero-Shot Reinforcement Learning Under Partial Observability
Behavior foundation models with GRU memory outperform memory-free zero-shot RL baselines in most partially observable ExORL settings, but the advantage is inconsistent on Cheetah.
Discussion (0). Continue with ORCID to comment.