Pith. sign in

Paper Citation Record · LEDGER

Is Exploration or Optimization the Problem for Deep Reinforcement Learning?

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2508.01329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01329 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:44:15.289881Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1eba9a00-639c-4083-85cb-f978864228db · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Towards Characterizing Divergence in Deep Q-Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:11.402366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:11.402366Z digest=sha256:ca5a6adaf5efcac3decc9072cfe6cf71985ace340e5d6fc33b2d6e5ad0230c86

Observation 0bd3b8ea-ca1a-401b-b5c1-521e33bd46b9 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning at the edge of the statistical precipice

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.445434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.523139Z digest=sha256:f0771b95b146622a521de291ecbc3b7c1d7e03beed9b2fe52a8648bb792d6076

Observation 9012bf15-39fb-4c85-9e88-ae2930c1ae75 · outbound

This paper cites Atari-5: Distilling the Arcade Learning Environment down to Five Games.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Atari-5: Distilling the Arcade Learning Environment down to Five Games

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:11.701517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:11.701517Z digest=sha256:29377488c008cf974b249b37f60337804dd90b4671043ec4c5f03e151d6de0d7

Observation 65cb8308-9a86-4ddb-991d-a34a13b634e2 · outbound

This paper cites Never give up: Learning directed exploration strategies, 2020.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Never give up: Learning directed exploration strategies, 2020

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.429262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.835102Z digest=sha256:46ea67cf46681a2618c76b0b6959b4308a4fc26bf3a0b0d596bf66b4ed21eb84

Observation c158a74f-ce46-4a51-ac71-a3108eda206b · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unifying count-based exploration and intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.414088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:11.997083Z digest=sha256:83f1ce1bfb2966b18a336dc052e5c965393bbb3764a14e4332ae951113031c1f

Observation e4b604a0-5918-4310-bf80-63c9e5ed83a8 · outbound

This paper cites Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T05:44:15.608184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.117380Z digest=sha256:cea891e432b874510954888e8761cd40eac7559461c5ceaba9239efa50f0fe3a

Observation 714980d0-1038-4f32-8146-408f74c02da7 · outbound

This paper cites Bellemare, Will Dabney, and R \' e mi Munos.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Will Dabney, and R \' e mi Munos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.399301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.284037Z digest=sha256:5fa0404e13d2cd1023169aabd9991cb5e3c64c85e1b843bbbb711c28e10cdff6

Observation 91486c0e-7de8-4ad9-9868-1988227725aa · outbound

This paper cites The theory of dynamic programming.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The theory of dynamic programming

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.384753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.418889Z digest=sha256:d5127224643153b495a49e2cb7ebb23bd50e15edeede8d2a1e4a9a2d297af93a

Observation 4caee5a4-d6dc-4201-a34c-10df449dec34 · outbound

This paper cites Interference and generalization in temporal difference learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Interference and generalization in temporal difference learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.369073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.518429Z digest=sha256:1992c277e019fab1d32a664dd2960188193f2d08af162ed67800629068aa35c4

Observation ab5064c3-2b99-465c-9b1d-573ff2ac0daf · outbound

This paper cites Exploration by random network distillation, 2018 a.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation, 2018 a

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.353804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.624275Z digest=sha256:09721802d2adb287d54a44537094a27bd823247ca0bda2cca8a03a3b6d16165b

Observation d30659ea-1eab-48c2-b419-4031954f0126 · outbound

This paper cites Exploration by random network distillation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Exploration by random network distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.339315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.800901Z digest=sha256:9bacba8a26b4dd5cf67d875dfc1d1466119cf35d6b6d621dc38da5b00531c28d

Observation 8f07d446-0420-4b67-8e80-e72e1a969202 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.323706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:12.940884Z digest=sha256:2b5449d4ed0e5a655aa8bfead6e1d7d5037793e7ae6fb4f5ecf4152e6c709c70

Observation 03652976-be24-49f7-bf0e-8ecf3ab81883 · outbound

This paper cites Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Target Network and Truncation Overcome The Deadly Triad in $Q$-Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:44:15.474846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.114051Z digest=sha256:3232088ed6eb48fc0fb3cbffa933947a8afad7668646358785cc239deb739a55

Observation b52edd0c-0814-40e9-be24-7cde3cdbe10e · outbound

This paper cites Phasic policy gradient.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Phasic policy gradient

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.308603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.226971Z digest=sha256:09f2458e39462f0374ff9723ee61ec113e7cec1942dc2c4938f697ac63c76428

Observation 88de9485-13b4-4789-8f86-17fa6b062881 · outbound

This paper cites Loss of plasticity in deep continual learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Loss of plasticity in deep continual learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.291200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.362972Z digest=sha256:b500f215ba6d38ddb609138de2ca6dcb71e3860a372fa201a3ed745a0b730e70

Observation 5a1084c3-0b78-49f7-b3a2-86d547419461 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep RL.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Stop regressing: Training value functions via classification for scalable deep RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.273604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.464882Z digest=sha256:e1bb7d1d1c3b697b2d205d8d0829dfae8830b4dcdc7a52c17e184ab7a55b0179

Observation 5ec7e74d-589f-4bf7-8881-693668b66d7b · outbound

This paper cites Fujimoto, H.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fujimoto, H

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.257618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.647048Z digest=sha256:504238bd304500d61c7387f593091795ba407194ea22003cc9fe9612a0f6db69

Observation 40ed1020-c78a-49f5-b408-d40e0818c4a5 · outbound

This paper cites Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:44:15.452859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.757638Z digest=sha256:620d2e2cd765cad8f98d354fd62a834f419f9333b8194ce0fe5b621e984e75e9

Observation 0a9231d9-a5f3-4db5-9681-80aabfe18382 · outbound

This paper cites Improving performance in reinforcement learning by breaking generalization in neural networks.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving performance in reinforcement learning by breaking generalization in neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.242508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:13.931406Z digest=sha256:211163ed3945895cd1b07d4fb00029f6dc9b6b48348ade606dc8bc8dd543082d

Observation 398644d4-ff48-4007-b3e2-65ca964f6c89 · outbound

This paper cites Haarnoja, A.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Haarnoja, A

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.224251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.108446Z digest=sha256:34965590f56b953a3cb640d42499c813675c6070b937302f057917d7fec6a869

Observation 3c9e5f52-3bf3-4823-b7ec-01384df6b968 · outbound

This paper cites Temporal difference learning for model predictive control.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning for model predictive control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.208501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.243178Z digest=sha256:25a3b7b052e854e035c15b21404d4d60c53ee805bac6e1c1f6f1e44a235d3674

Observation db07826b-be36-4c5a-baa9-4954c7e2210d · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Rainbow: Combining improvements in deep reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.191803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.406629Z digest=sha256:08b8e8672b9cd70eb384557888ae48e83dae2f7fa47d8363ae76fc55e668d182

Observation db39975b-1a29-41ba-870f-2a6b4559fffe · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Approximately optimal approximate reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.174366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.537137Z digest=sha256:0c9e1765800f8f77ce3c7f187668f5d143cb571bfc27f544f23555d6508b3744

Observation 8e9f4d0f-9c87-423d-80bf-f327d0e7ff56 · outbound

This paper cites Discor: Corrective feedback in reinforcement learning via distribution correction.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Discor: Corrective feedback in reinforcement learning via distribution correction

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.158893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.662962Z digest=sha256:0412aec796d73fca249a942fcdfec1a0a0da3d8513ffdd17b45389596f626844

Observation 86f10e3c-19d3-43be-831d-836dc2b3b1bb · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.141186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.793980Z digest=sha256:56daeb4c32629e037022235ed6372a77ca322af1b76cce03d60588983985bf6a

Observation b9d82107-e6cb-4913-ae78-48ec50239c3b · outbound

This paper cites Maxmin q-learning: Controlling the estimation bias of q-learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Maxmin q-learning: Controlling the estimation bias of q-learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.125073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:14.974012Z digest=sha256:82b6c1c65757a078693fd30dc694430a8dbec10c34e8bd98ee1c3d3696c7eb98

Observation 0d0b4467-759a-436b-bcc0-6d29e243fd3c · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.108519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.071409Z digest=sha256:0a32a5a3a1112a27dea954cca945f6b1360af9af32d0f17d383ec18206653ffa

Observation a9e2609a-7fab-4c54-b406-4663a7ae6527 · outbound

This paper cites Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Hyar: Addressing discrete-continuous action reinforcement learning via hybrid action representation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.092348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.090226Z digest=sha256:70ec61817ba1fdb842c0a9f06452031a90243d7b225db6c0433199c9e619dc43

Observation 02bf594d-fa09-4e4c-89fa-1bd7856647a7 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:16.076826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.102829Z digest=sha256:b8b0ce67c43f3d346e196a13215abc833f0959746a27658c3a2be14c7c58a6eb

Observation 692c1348-e985-40f7-b5ff-a4ec1e411a5d · outbound

This paper cites Understanding plasticity in neural networks.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Understanding plasticity in neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.060939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.107894Z digest=sha256:3ae21f8dc3a35ee88f35556f819dea22aaf12f764361ccff14dd41578c29711b

Observation 296b198c-6a76-4466-a25c-824730fbc677 · outbound

This paper cites Normalization and effective learning rates in reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Normalization and effective learning rates in reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.112426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.112426Z digest=sha256:aed91635fec742a94190aaa485e2dae4b1b8b612aed7472370246f15864318d0

Observation 5c2f20bc-4c50-40a8-bf90-2a42869a932b · outbound

This paper cites Human-level control through deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Human-level control through deep reinforcement learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.117368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.117368Z digest=sha256:d3933c799d440675879a522685a1b3ab7483748ec1b1084cafcff1036ad6d063

Observation 8a11509c-8cd1-46bb-8bef-8b133a160193 · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.122140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.122140Z digest=sha256:6c6e62f1f6bc311fb84239f6af47d5cee80c7025025f04f770beca704b256202

Observation 64d31f03-dd55-440d-a529-5d9f1229d620 · outbound

This paper cites Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.127132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.127132Z digest=sha256:37eaa75fba6682ffb6f6a14e5be69583f2afe774389234adc3f35c5459b235ec

Observation 73c79c66-b119-417c-b4b2-d90458ca0b5e · outbound

This paper cites Nikishin, M.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Nikishin, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.032454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.132767Z digest=sha256:1b692aa7d4721bfe2b7f565f3fe0473281c1ec83dd658ef50af9d252c142be9b

Observation 49e4a198-8e23-4210-a2a6-6a14c665cb30 · outbound

This paper cites Mixtures of Experts Unlock Parameter Scaling for Deep RL.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Mixtures of Experts Unlock Parameter Scaling for Deep RL

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.137915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.137915Z digest=sha256:2170b14a5508880040ac98628183e8bfb430cc1bff039d3289d0c6d177a42eaa

Observation befae0ed-fecd-4bb2-92b0-127d5f9bb8d5 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, 2022.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Chatgpt: Optimizing language models for dialogue, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.016932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.142696Z digest=sha256:5f610ffa9db04557329f5b6d34cd04ee6cf5bbdc38b28970f45fece61f0a5162

Observation 83f468a9-7d71-4521-bab2-62eb0ccd4f4b · outbound

This paper cites Dota 2 with large scale deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Dota 2 with large scale deep reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:16.002008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.147758Z digest=sha256:51c2503ea92d9f071fabe451f59279b549ebc4d6576d2fb7d9478c6a2db3a64e

Observation 6c23f6f7-da7b-4a29-bff5-dd379dde39c9 · outbound

This paper cites Bellemare, Aaron van den Oord, and Remi Munos.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bellemare, Aaron van den Oord, and Remi Munos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.987483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.152329Z digest=sha256:f3aa84ddd0210f14ca90c1132fd1f8b5767a8867c99b9b509b43c9c5c7060dbc

Observation 042f1dbe-316a-4063-aabf-d0de3c46f816 · outbound

This paper cites The difficulty of passive learning in deep reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? The difficulty of passive learning in deep reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.972739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.157046Z digest=sha256:d4a99b166eb0ef619333b180a8779195e452710cc13f25d7b1b5bf8fc98a7eb0

Observation d0ae13c7-9db3-427e-84c4-660e1cd560dd · outbound

This paper cites Jha, Toshisada Mariyama, and Daniel Nikovski.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Jha, Toshisada Mariyama, and Daniel Nikovski

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.956169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.161903Z digest=sha256:e89ec0696c2115559fc75335ba29c8b5c773c6b7e81bf53cf6d5de99830bd5bc

Observation 34226b10-1c78-4dad-a4a4-e40eae71538b · outbound

This paper cites Fuzzy tiling activations: A simple approach to learning sparse representations online.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Fuzzy tiling activations: A simple approach to learning sparse representations online

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.940986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.167501Z digest=sha256:e62581fa20732694ba9ffc83bd1446b18adcdad8be0275d0b5c51b15f2e59a88

Observation eb97cefc-621b-43d2-a46e-78438b04dc6c · outbound

This paper cites Efros, and Trevor Darrell.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Efros, and Trevor Darrell

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.926529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.172719Z digest=sha256:11e67a7ea25b4af8f093b9d9ad15137d13b4e7d45931929f29a8115a41b84604

Observation 3f8cf745-15fe-495c-8866-5c74e443db44 · outbound

This paper cites Bridging the gap between target networks and functional regularization.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Bridging the gap between target networks and functional regularization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.909906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.179837Z digest=sha256:8b29296e746b630e98dcfd22b5422558969cb0daf945be8d7f177d9f3921832c

Observation d6832d5f-b75a-401a-b8b8-187ff913c706 · outbound

This paper cites Decoupling value and policy for generalization in reinforcement learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Decoupling value and policy for generalization in reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.895196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.185252Z digest=sha256:d4ac7eb8092ceab44f93c1b292617a68b6ed0f6103214e7629cc31f90a0405b1

Observation b7edd74f-59f8-414a-98b4-47d23bbb7de1 · outbound

This paper cites Prioritized experience replay.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Prioritized experience replay

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.880540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.190676Z digest=sha256:13c969b4a7fa102449e576587fdb0162d7b8fc6b3c944c286e45b26c7005b837

Observation 9c29176a-8c64-4fbb-8bf2-d94bd3197d30 · outbound

This paper cites Schulman, S.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Schulman, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.864276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.196969Z digest=sha256:1a7fc0dac179cd1af5b628f8fccead41dda37465d88069c1375ae6901fd07cc4

Observation baeb9d29-7246-4a99-ba62-f1f95bb9e9ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Proximal Policy Optimization Algorithms

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.201847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.201847Z digest=sha256:5078bc38d68192dee02a4c88b8cb885a925ead0685882884663cc337839ea63a

Observation 339911d0-db30-44f0-a4ee-5bfce67c664a · outbound

This paper cites Courville, Marc G.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Courville, Marc G

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.849518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.206930Z digest=sha256:f257a2af97363ec44a325361b4d93f94b73c14026b220b230640d0eb86a3f192

Observation 0fce935c-1f8c-41d6-b479-817a1160e7f8 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:15.834305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.212085Z digest=sha256:5384d1e0a75d1edd64d896eef2df5e4dcae4591e97604b6def3cd858271b8304

Observation f803f178-40de-4eb6-b7de-bdb65e7ad3b7 · outbound

This paper cites \#exploration: A study of count-based exploration for deep reinforcement learning, 2017.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? \#exploration: A study of count-based exploration for deep reinforcement learning, 2017

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.817834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.216775Z digest=sha256:e75c06543c730aeb7f6a6788fdf7ddcb91cdc40c3538858f9a8f1b891406300c

Observation b5bdd617-acf5-4402-9c79-250644d4e6ae · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.802566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.221980Z digest=sha256:ad35be7ffa4c797b50bb3d21d1f6af6dd2418a2ae65ae14a0786dff9db729e14

Observation b92963fa-2d7e-45e5-8748-23f0767b9695 · outbound

This paper cites Temporal difference learning and td-gammon.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Temporal difference learning and td-gammon

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.227056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.227056Z digest=sha256:743df7b317bcd8cb6d04700651866750900c7a151de10d4f3e926123bb846e4b

Observation 39dc3e1e-6ffe-45ae-b113-8b0222107181 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep reinforcement learning with double q-learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.232608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.232608Z digest=sha256:561379131c6cc9b0b74c486bcc1883e4d22f9f4539667e521375fd31f997f001

Observation 4d54ad17-5829-4f9c-a0f1-9b53f8d0d437 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Deep Reinforcement Learning and the Deadly Triad

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.237995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.237995Z digest=sha256:4c91f73b9fdcf002e8851a9662802ac14f87b79bcf6513529ee579b6b024bbdd

Observation 4711cce7-329c-4425-8ae4-aac846a25fca · outbound

This paper cites Overcoming the spectral bias of neural value approximation.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Overcoming the spectral bias of neural value approximation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.766412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.242757Z digest=sha256:123ef1449c3b459bfbed7fb0925734656bc4b97034192fa5f7daa6e06243081e

Observation ac19e42c-0955-4e59-9db1-4a4516f45894 · outbound

This paper cites MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? MinAtar : An atari-inspired testbed for thorough and reproducible reinforcement learning experiments

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.749648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.247375Z digest=sha256:f3a7808f8f1722499c496b865bd3d3076ebe1a526a9be9f68e8aa8fd25186a31

Observation 60d6ec9f-97b9-48af-8fe1-a0a208a49cc9 · outbound

This paper cites Learning invariant representations for reinforcement learning without reconstruction.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Learning invariant representations for reinforcement learning without reconstruction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.734634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.251844Z digest=sha256:e1a611fa08bb6181d2b50f01c13108ea57d8f2904c8236d54a7685464df16687

Observation fdbd5858-aa05-4113-a705-2a16f45e2761 · outbound

This paper cites Gendice: Generalized offline estimation of stationary values.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Gendice: Generalized offline estimation of stationary values

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.719434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.256785Z digest=sha256:220fd5b2244e33b3454573ba24b4dee12be95db793bd48f5fbf29844fa37a2e6

Observation 719dedf7-6b7d-4556-ad52-758912305210 · outbound

This paper cites Breaking the deadly triad with a target network.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Breaking the deadly triad with a target network

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.704127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.262635Z digest=sha256:da42aad3954b0ab13e2ebeef88993372635b4d9b4d677080bff39497f0225b95

Observation 76405fa5-0000-4990-93c4-51b2eff046fd · outbound

This paper cites BeBold: Exploration Beyond the Boundary of Explored Regions.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? BeBold: Exploration Beyond the Boundary of Explored Regions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.267730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.267730Z digest=sha256:82d37a82e1e4e70083a39cf288e91af3feb05fb54528fe1bac928a8bf3df1221

Observation d4bfed1c-e7c4-4ae2-8c29-abd91c3ee521 · outbound

This paper cites Noveld: A simple yet effective exploration criterion.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Noveld: A simple yet effective exploration criterion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:44:15.689132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.273651Z digest=sha256:16b15c423f1e06fa48e1c7a02c106a5655ac6df663e692e927d48ac4d444589e

Observation 3ac185bc-b5ee-42b2-a8b6-661b1dc03eea · outbound

This paper cites @esa (Ref.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.279905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.279905Z digest=sha256:9d5964621e3feac19e848a7793fd62f300c5c5d60c30b8cd3b68cea051a244f0

Observation 2de190b9-18af-4c9c-a78a-2fde7f7fe7d6 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:15.285200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:15.285200Z digest=sha256:5e07f8a14e3579cebb7c1cf411d8e0c8017fc5d81dec95600986b4b591cb7cd2

Observation b258ad46-e590-4175-8533-77a1ac54c371 · outbound

This paper cites an unresolved cited work.

Is Exploration or Optimization the Problem for Deep Reinforcement Learning? Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:44:15.653607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T05:44:15.289881Z digest=sha256:7b90276b8f3ffcc1c86dd6d79ae7b9d21f7c7b746d1eff63e19b68528fff8ee6

Pith citing papers

No inbound Pith citation observations are available.