Temporal Difference Uncertainties as a Signal for Exploration

Alexandre Galashov; Andre Barreto; Diana L. Borsa; Francesco Visin; Jane X. Wang; Nicolas Heess; Pablo Sprechmann; Razvan Pascanu; Sebastian Flennerhag; Steven Kapturowski

arxiv: 2010.02255 · v2 · pith:NT75H6IGnew · submitted 2020-10-05 · 💻 cs.AI · cs.LG· stat.ML

Temporal Difference Uncertainties as a Signal for Exploration

Sebastian Flennerhag , Jane X. Wang , Pablo Sprechmann , Francesco Visin , Alexandre Galashov , Steven Kapturowski , Diana L. Borsa , Nicolas Heess

show 2 more authors

Andre Barreto Razvan Pascanu

This is my paper

classification 💻 cs.AI cs.LGstat.ML

keywords explorationuncertaintyvalueagentdifferenceestimateslearningtemporal

0 comments

read the original abstract

An effective approach to exploration in reinforcement learning is to rely on an agent's uncertainty over the optimal policy, which can yield near-optimal exploration strategies in tabular settings. However, in non-tabular settings that involve function approximators, obtaining accurate uncertainty estimates is almost as challenging a problem. In this paper, we highlight that value estimates are easily biased and temporally inconsistent. In light of this, we propose a novel method for estimating uncertainty over the value function that relies on inducing a distribution over temporal difference errors. This exploration signal controls for state-action transitions so as to isolate uncertainty in value that is due to uncertainty over the agent's parameters. Because our measure of uncertainty conditions on state-action transitions, we cannot act on this measure directly. Instead, we incorporate it as an intrinsic reward and treat exploration as a separate learning problem, induced by the agent's temporal difference uncertainties. We introduce a distinct exploration policy that learns to collect data with high estimated uncertainty, which gives rise to a curriculum that smoothly changes throughout learning and vanishes in the limit of perfect value estimates. We evaluate our method on hard exploration tasks, including Deep Sea and Atari 2600 environments and find that our proposed form of exploration facilitates both diverse and deep exploration.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Subjective functions
cs.AI 2025-12 unverdicted novelty 6.0

Subjective functions are proposed as endogenous higher-order objectives, with expected prediction error as a concrete example, to explain goal synthesis in intelligent agents.