Pith. sign in

Paper Citation Record · LEDGER

Is Exploration or Optimization the Problem for Deep Reinforcement Learning?

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.01329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01329 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:44:15.289881Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1eba9a00-639c-4083-85cb-f978864228db · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Towards Characterizing Divergence in Deep Q-Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:11.402366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:11.402366Z digest=sha256:03a407196601213e761b408f3bf7827cedb43f3b6083f11b21e8706003212b92

Observation 0bd3b8ea-ca1a-401b-b5c1-521e33bd46b9 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning at the edge of the statistical precipice

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.445434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.523139Z digest=sha256:976f0901ff80ad6d13caf269c299d4818cd898183d74e11fe7de817e0708ad38

Observation 9012bf15-39fb-4c85-9e88-ae2930c1ae75 · outbound

This paper cites Atari-5: Distilling the Arcade Learning Environment down to Five Games.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Atari-5: Distilling the Arcade Learning Environment down to Five Games

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:11.701517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:11.701517Z digest=sha256:b0c09e24614448b8fbed0ddee49ceef5d090a8137a87f07b483176225d4b3f34

Observation 65cb8308-9a86-4ddb-991d-a34a13b634e2 · outbound

This paper cites Never give up: Learning directed exploration strategies, 2020.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Never give up: Learning directed exploration strategies, 2020

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.429262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.835102Z digest=sha256:e5bf2373a2dc3bae6382ff8fe4ebf0be375dc2c4292239749564e522c65faf21

Observation c158a74f-ce46-4a51-ac71-a3108eda206b · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unifying count-based exploration and intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.414088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.997083Z digest=sha256:6242a77539974abbea79b82956a7a080f4b2bf5a0d757380a02d188e2a69b711

Observation e4b604a0-5918-4310-bf80-63c9e5ed83a8 · outbound

This paper cites Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T05:44:15.608184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.117380Z digest=sha256:a32ef4b585375340c727e3439431fd756fb1f752f106e77acbc09b5da203f0c0

Observation 714980d0-1038-4f32-8146-408f74c02da7 · outbound

This paper cites Bellemare, Will Dabney, and R \' e mi Munos.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Will Dabney, and R \' e mi Munos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.399301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.284037Z digest=sha256:24dfc59970389de19e40d04324e1b635ae6c0f8faa8c6d8911823389ff64e856

Observation 91486c0e-7de8-4ad9-9868-1988227725aa · outbound

This paper cites The theory of dynamic programming.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The theory of dynamic programming

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.384753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.418889Z digest=sha256:7a11844e945e0b75073ee441b62737cdc125bf5d83234e1c82e91872c7739a4e

Observation 4caee5a4-d6dc-4201-a34c-10df449dec34 · outbound

This paper cites Interference and generalization in temporal difference learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Interference and generalization in temporal difference learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.369073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.518429Z digest=sha256:49e03ade6cc3d68df31381c0ebc1baa4718e6fdf56c2ce8a46ce10a99d56aa32

Observation ab5064c3-2b99-465c-9b1d-573ff2ac0daf · outbound

This paper cites Exploration by random network distillation, 2018 a.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation, 2018 a

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.353804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.624275Z digest=sha256:9575a2e262f3055e9f0fe3d063d111a94f2a2e5d2ec9a0c399d688d7d1fd58e5

Observation d30659ea-1eab-48c2-b419-4031954f0126 · outbound

This paper cites Exploration by random network distillation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.339315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.800901Z digest=sha256:7bf3dcdb821ad462e8e97da07b785db533bd816f5add1f6719030a7050a75488

Observation 8f07d446-0420-4b67-8e80-e72e1a969202 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.323706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.940884Z digest=sha256:dc80b6cc3b6ddf755f5344660e89ae2d71a22b1f4aeacd038a53058c2800dd77

Observation 03652976-be24-49f7-bf0e-8ecf3ab81883 · outbound

This paper cites Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:44:15.474846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.114051Z digest=sha256:1159c3635e9271c2b34c130019ff951efaf1a999b9d3f50c0ea74774f83f6bdd

Observation b52edd0c-0814-40e9-be24-7cde3cdbe10e · outbound

This paper cites Phasic policy gradient.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Phasic policy gradient

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.308603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.226971Z digest=sha256:3d4c0d78f3a235f18f64ca2497791937de383e2dde549954a4b95a9878749646

Observation 88de9485-13b4-4789-8f86-17fa6b062881 · outbound

This paper cites Loss of plasticity in deep continual learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Loss of plasticity in deep continual learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.291200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.362972Z digest=sha256:49243e4e38d80f5a2a3a7977d41073280437a05dcea7b40466ae5f727d1a8176

Observation 5a1084c3-0b78-49f7-b3a2-86d547419461 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep RL.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Stop regressing: Training value functions via classification for scalable deep RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.273604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.464882Z digest=sha256:6acf15aea429d41ba4b19af8b4b19312df9b7974065dbeef8727b3719ed4b8db

Observation 5ec7e74d-589f-4bf7-8881-693668b66d7b · outbound

This paper cites Fujimoto, H.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fujimoto, H

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.257618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.647048Z digest=sha256:ef079b9cc03debc942b31a863e29bd9e1e642ed078d1bdf700cc5115dec1f3ec

Observation 40ed1020-c78a-49f5-b408-d40e0818c4a5 · outbound

This paper cites Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:44:15.452859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.757638Z digest=sha256:b6fa92dd4c5e96f52c0672a9c7b5c784a27ad4f355de4d625d187728379612a9

Observation 0a9231d9-a5f3-4db5-9681-80aabfe18382 · outbound

This paper cites Improving performance in reinforcement learning by breaking generalization in neural networks.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving performance in reinforcement learning by breaking generalization in neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.242508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.931406Z digest=sha256:a1a283f0ddfca51ac491be145379f472f1d572f854e2f4fe00d857a951e02b14

Observation 398644d4-ff48-4007-b3e2-65ca964f6c89 · outbound

This paper cites Haarnoja, A.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Haarnoja, A

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.224251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.108446Z digest=sha256:df83b178daef1214c6d19283ada220a10a8809a4e1c39fbbf2163795c22412a7

Observation 3c9e5f52-3bf3-4823-b7ec-01384df6b968 · outbound

This paper cites Temporal difference learning for model predictive control.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning for model predictive control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.208501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.243178Z digest=sha256:af239de7ac7f63f4aa5bd575af0b17a686f256092f5d0fb7027becce8f2aa84c

Observation db07826b-be36-4c5a-baa9-4954c7e2210d · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Rainbow: Combining improvements in deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.191803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.406629Z digest=sha256:61db466f9e40ef13fc4f604535380dd5c9a6bd1b8e4a2e9b22d37e3beac7abb1

Observation db39975b-1a29-41ba-870f-2a6b4559fffe · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Approximately optimal approximate reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.174366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.537137Z digest=sha256:92382edebcb2abbbdce7792487d28a69ca9afa24616dd438f153fa84b863928d

Observation 8e9f4d0f-9c87-423d-80bf-f327d0e7ff56 · outbound

This paper cites Discor: Corrective feedback in reinforcement learning via distribution correction.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Discor: Corrective feedback in reinforcement learning via distribution correction

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.158893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.662962Z digest=sha256:693a08d12e79b9d76bd9e26d3bd75ad5c7aba76b4f8777e3a4b05271c46dcef0

Observation 86f10e3c-19d3-43be-831d-836dc2b3b1bb · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.141186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.793980Z digest=sha256:a1bd89ee34944fe39f2130b66a6da414505e82b3e360619f2e4ed8a49d0e8a42

Observation b9d82107-e6cb-4913-ae78-48ec50239c3b · outbound

This paper cites Maxmin q-learning: Controlling the estimation bias of q-learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Maxmin q-learning: Controlling the estimation bias of q-learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.125073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.974012Z digest=sha256:d04fc32d2ad334a531df430c5976bf866d2c01710752cea23933c825ce17cc5d

Observation 0d0b4467-759a-436b-bcc0-6d29e243fd3c · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.108519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.071409Z digest=sha256:35e32110631360d9703f9f8b07bb1e90db5e73c1db967240ca0c8ab7393ca359

Observation a9e2609a-7fab-4c54-b406-4663a7ae6527 · outbound

This paper cites Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.092348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.090226Z digest=sha256:732907b7845ef9510f5f9a2a7f5f7689612bfeabfbe82bc583a15c3c68fcfd48

Observation 02bf594d-fa09-4e4c-89fa-1bd7856647a7 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.076826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.102829Z digest=sha256:7712de61429bd32735a75e79e47b960319d9214a89dcfcb878dbcb57e0417cb2

Observation 692c1348-e985-40f7-b5ff-a4ec1e411a5d · outbound

This paper cites Understanding plasticity in neural networks.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Understanding plasticity in neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.060939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.107894Z digest=sha256:33ef75ddd9832215a3aa1588640936a407469f042bace446a822bc9dcf4be225

Observation 296b198c-6a76-4466-a25c-824730fbc677 · outbound

This paper cites Normalization and effective learning rates in reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Normalization and effective learning rates in reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.112426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.112426Z digest=sha256:99465dc9d6b162899a032e456f07c791b0344b842526f3e81e0925b91538bd9c

Observation 5c2f20bc-4c50-40a8-bf90-2a42869a932b · outbound

This paper cites Human-level control through deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Human-level control through deep reinforcement learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.117368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.117368Z digest=sha256:1805c4fb184b97de1129a44cdb7f86e5bd1cca9be9287f9abe6e31eca451b281

Observation 8a11509c-8cd1-46bb-8bef-8b133a160193 · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.122140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.122140Z digest=sha256:17c4c9bdc74cc416033ec2cdb7af24e7369c120c9586fe805a441797141ea620

Observation 64d31f03-dd55-440d-a529-5d9f1229d620 · outbound

This paper cites Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.127132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.127132Z digest=sha256:0dfa39cdc3837060f0113cb91167a2d12e5858a023f28c224cd31359f02a4cc9

Observation 73c79c66-b119-417c-b4b2-d90458ca0b5e · outbound

This paper cites Nikishin, M.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Nikishin, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.032454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.132767Z digest=sha256:2e0f8ee40015b5bf026e39508501393ec933d1cb41e67397e16c47849d8d1f33

Observation 49e4a198-8e23-4210-a2a6-6a14c665cb30 · outbound

This paper cites Mixtures of Experts Unlock Parameter Scaling for Deep RL.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Mixtures of Experts Unlock Parameter Scaling for Deep RL

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.137915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.137915Z digest=sha256:35c98f4c9a12ffcf8764443258381c027a0aecfa8e5f004b460320d2da3a8faa

Observation befae0ed-fecd-4bb2-92b0-127d5f9bb8d5 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, 2022.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Chatgpt: Optimizing language models for dialogue, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.016932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.142696Z digest=sha256:f55aa342837b2fa7dd67d3c8ef737c065e13cf2b47251830061aed3911724f3e

Observation 83f468a9-7d71-4521-bab2-62eb0ccd4f4b · outbound

This paper cites Dota 2 with large scale deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Dota 2 with large scale deep reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.002008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.147758Z digest=sha256:59a6edf9c1736055c499a5344f97802045706c2717a50d2b5bda7b5d380e1168

Observation 6c23f6f7-da7b-4a29-bff5-dd379dde39c9 · outbound

This paper cites Bellemare, Aaron van den Oord, and Remi Munos.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Aaron van den Oord, and Remi Munos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.987483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.152329Z digest=sha256:0a30d92b54ded266a705907710068d45aaffc55e51953d998c94f0f16af7a497

Observation 042f1dbe-316a-4063-aabf-d0de3c46f816 · outbound

This paper cites The difficulty of passive learning in deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The difficulty of passive learning in deep reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.972739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.157046Z digest=sha256:a7250d7cca162a627fac315ec5b0253721dc52e9bc88e92da18e51edc60dde22

Observation d0ae13c7-9db3-427e-84c4-660e1cd560dd · outbound

This paper cites Jha, Toshisada Mariyama, and Daniel Nikovski.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Jha, Toshisada Mariyama, and Daniel Nikovski

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.956169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.161903Z digest=sha256:1f1ff091665ad90beba9d08052af4ea5df476bb2db55439aa1f6a08b0ae3402c

Observation 34226b10-1c78-4dad-a4a4-e40eae71538b · outbound

This paper cites Fuzzy tiling activations: A simple approach to learning sparse representations online.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fuzzy tiling activations: A simple approach to learning sparse representations online

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.940986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.167501Z digest=sha256:0826c0dd3542fff17927f102fae57004b11df5ec0924da306bd5adba76c9c3c2

Observation eb97cefc-621b-43d2-a46e-78438b04dc6c · outbound

This paper cites Efros, and Trevor Darrell.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Efros, and Trevor Darrell

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.926529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.172719Z digest=sha256:5cab9cd0842963914ed4b627a6528fbda7483bb5fffccba8f87b807d5a6527c0

Observation 3f8cf745-15fe-495c-8866-5c74e443db44 · outbound

This paper cites Bridging the gap between target networks and functional regularization.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bridging the gap between target networks and functional regularization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.909906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.179837Z digest=sha256:bab92ea6bfa5be5bf97ad307dd16ca7d06853617ec0caaa021f218da8754a9a0

Observation d6832d5f-b75a-401a-b8b8-187ff913c706 · outbound

This paper cites Decoupling value and policy for generalization in reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Decoupling value and policy for generalization in reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.895196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.185252Z digest=sha256:9bc343707a88d39be1171a5435ecee30a5755da55c5ec4dc917dc78b7fa1dc6c

Observation b7edd74f-59f8-414a-98b4-47d23bbb7de1 · outbound

This paper cites Prioritized experience replay.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Prioritized experience replay

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.880540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.190676Z digest=sha256:d8bf86d659f59b2c5569323e48769d3486b011100a1ba69a5f069c934aea9804

Observation 9c29176a-8c64-4fbb-8bf2-d94bd3197d30 · outbound

This paper cites Schulman, S.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Schulman, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.864276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.196969Z digest=sha256:8c5219c3dd500f8397e68ec33a9fa6b2f4fe70d9179821298f5c145d4eb7d869

Observation baeb9d29-7246-4a99-ba62-f1f95bb9e9ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Proximal Policy Optimization Algorithms

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.201847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.201847Z digest=sha256:982e732955879ba72f3efc6a74e5d94b4993bb4aaa729014dcb6f1bc55d86319

Observation 339911d0-db30-44f0-a4ee-5bfce67c664a · outbound

This paper cites Courville, Marc G.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Courville, Marc G

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.849518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.206930Z digest=sha256:54eeb9a1b0e1b4b607e18e497ae2363b44496134463ffdf84620f54ec8c56a20

Observation 0fce935c-1f8c-41d6-b479-817a1160e7f8 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:15.834305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.212085Z digest=sha256:4474abca727daecce4de77e94d2918fb34025e7e5b81f279985accda060c5106

Observation f803f178-40de-4eb6-b7de-bdb65e7ad3b7 · outbound

This paper cites \#exploration: A study of count-based exploration for deep reinforcement learning, 2017.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? \#exploration: A study of count-based exploration for deep reinforcement learning, 2017

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.817834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.216775Z digest=sha256:04daeb0ef08472ff0a2d31676b2bbcbb749b47b146836cc5d34626be6054272e

Observation b5bdd617-acf5-4402-9c79-250644d4e6ae · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.802566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.221980Z digest=sha256:4af4c88cbba06b26cbb6f221d17cc4fd001a55acf9713c49519100d68fee3088

Observation b92963fa-2d7e-45e5-8748-23f0767b9695 · outbound

This paper cites Temporal difference learning and td-gammon.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning and td-gammon

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.227056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.227056Z digest=sha256:b287718779fbbf33d9c9adf3723cd2d9ebba1cd540360137b70412124fc73141

Observation 39dc3e1e-6ffe-45ae-b113-8b0222107181 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning with double q-learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.232608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.232608Z digest=sha256:ac28476fd3cae9b7dfe470550328522aa87693eba10be2b17a56957d36b0cf98

Observation 4d54ad17-5829-4f9c-a0f1-9b53f8d0d437 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep Reinforcement Learning and the Deadly Triad

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.237995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.237995Z digest=sha256:4637449c7787b68444d69f2b2640b518dc08d7b8a7cb1b94a7e634ea9670f757

Observation 4711cce7-329c-4425-8ae4-aac846a25fca · outbound

This paper cites Overcoming the spectral bias of neural value approximation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Overcoming the spectral bias of neural value approximation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.766412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.242757Z digest=sha256:3a3f7ea7669bbe5b71220fbf3641b21a591a959a664ec69eafaa101cadf59d97

Observation ac19e42c-0955-4e59-9db1-4a4516f45894 · outbound

This paper cites MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.749648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.247375Z digest=sha256:097c7e0fcc10fe70e671c6155cc903de56c154e02d1bb86ad6da8b03130e716d

Observation 60d6ec9f-97b9-48af-8fe1-a0a208a49cc9 · outbound

This paper cites Learning invariant representations for reinforcement learning without reconstruction.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Learning invariant representations for reinforcement learning without reconstruction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.734634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.251844Z digest=sha256:69e2d020b518e5bfe8618868a249af125a56def437333613776e46fb5f3c4904

Observation fdbd5858-aa05-4113-a705-2a16f45e2761 · outbound

This paper cites Gendice: Generalized offline estimation of stationary values.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Gendice: Generalized offline estimation of stationary values

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.719434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.256785Z digest=sha256:0fc9889e3a32940027c561082a46ba226c8178866cb40232df92ae56a7fe8c6f

Observation 719dedf7-6b7d-4556-ad52-758912305210 · outbound

This paper cites Breaking the deadly triad with a target network.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Breaking the deadly triad with a target network

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.704127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.262635Z digest=sha256:8b320b5717e0d663e346cc3199c623d976e9ea3991b86a26ae2bbebabdddce96

Observation 76405fa5-0000-4990-93c4-51b2eff046fd · outbound

This paper cites BeBold: Exploration Beyond the Boundary of Explored Regions.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? BeBold: Exploration Beyond the Boundary of Explored Regions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.267730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.267730Z digest=sha256:18cd05d3c8773062e1f869d62aab21c747d176f4df439dc4a32951f43faa6c2f

Observation d4bfed1c-e7c4-4ae2-8c29-abd91c3ee521 · outbound

This paper cites Noveld: A simple yet effective exploration criterion.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Noveld: A simple yet effective exploration criterion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.689132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.273651Z digest=sha256:54c53c73c4b20931c80b24a3714dbc0df23b70445d1d3d8330f50b1f094653bd

Observation 3ac185bc-b5ee-42b2-a8b6-661b1dc03eea · outbound

This paper cites @esa (Ref.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.279905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.279905Z digest=sha256:6a88587599be2244a24e6cd6051dbc5c52d639a78acb77ef500e65c9accce6d5

Observation 2de190b9-18af-4c9c-a78a-2fde7f7fe7d6 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.285200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.285200Z digest=sha256:93dca1d8e8db51c0f9f97e99dd5d6b457ccd3ce915692e71a4eab498ecd81dfa

Observation b258ad46-e590-4175-8533-77a1ac54c371 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:15.653607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.289881Z digest=sha256:41e981ab56b2b8256deb6fd07608097c60f08481b8a7b72dd03529a526d15350

Pith citing papers

No inbound Pith citation observations are available.