Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:12.221295Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.11044.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:12.221295Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:59.951972Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:08:00.455034Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0eb65fc1-9d28-4857-9db5-a9974d1e5a56 · outbound
Exploration by Random Distribution Distillation Deep exploration via bootstrapped dqn,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f016da28-af73-4286-b8e6-a27bb6062496 · outbound
Exploration by Random Distribution Distillation An analysis of model-based interval estimation for markov decision processes,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61e4d41-e894-48fd-b0ab-a672e6ba11ad · outbound
Exploration by Random Distribution Distillation Minimax regret bounds for reinforcement learning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9725abb0-e729-4bfd-8f64-7eeb29a7b53b · outbound
Exploration by Random Distribution Distillation The Alberta Plan for AI Research
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3b377c-e0b8-4b17-9c1a-0d476d9926ce · outbound
Exploration by Random Distribution Distillation Unifying count-based exploration and intrinsic motivation,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 817e8a24-ba9f-49fe-8186-7c05d023cdf9 · outbound
Exploration by Random Distribution Distillation Flipping coins to estimate pseudocounts for exploration in reinforcement learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 516956b7-2f51-493f-8ad2-2d9ce29e0446 · outbound
Exploration by Random Distribution Distillation Count-based exploration with neural density models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8eb97b88-0ec9-4cf7-8bfb-4439e6cb9798 · outbound
Exploration by Random Distribution Distillation Count-based exploration with the successor representation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ca4efab-9682-45c3-98d6-2f9852e4cf09 · outbound
Exploration by Random Distribution Distillation Curiosity-driven exploration by self- supervised prediction,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6877a50-f0f6-43a4-9779-cc36f7fa25b5 · outbound
Exploration by Random Distribution Distillation Exploration by Random Network Distillation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aaf07ee-afad-4a18-bc13-95d8f3f07399 · outbound
Exploration by Random Distribution Distillation Is q-learning provably efficient?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8115b221-7f28-4551-9109-4bb4ae93f11b · outbound
Exploration by Random Distribution Distillation Exploration and anti-exploration with distributional random network distillation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6615c679-dfe5-454b-a919-25d1a598fbb3 · outbound
Exploration by Random Distribution Distillation Q-learning,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc62367-2383-4049-8f92-ef7e0c2e6340 · outbound
Exploration by Random Distribution Distillation Human-level control through deep rein- forcement learning,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94a3152a-ea97-4944-85ae-e239488fc964 · outbound
Exploration by Random Distribution Distillation Rainbow: Combining improvements in deep reinforcement learning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c4edd3f9-b9a4-4f3f-b4b9-0ac936755f1d · outbound
Exploration by Random Distribution Distillation Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fd5aa7e3-bdec-4dca-b5e0-599adafbf90c · outbound
Exploration by Random Distribution Distillation Using confidence bounds for exploitation-exploration trade-offs,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 71522f23-58ac-4361-9bb3-4d77e6c40594 · outbound
Exploration by Random Distribution Distillation Improving generalization for temporal difference learning: The successor representa- tion,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97705774-3838-4cb9-b896-fc0e311ec1b2 · outbound
Exploration by Random Distribution Distillation On Bonus-Based Exploration Methods in the Arcade Learning Environment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5734298d-bb8e-4b35-b5f2-1a368f586323 · outbound
Exploration by Random Distribution Distillation # exploration: A study of count-based exploration for deep reinforcement learning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 59e9d073-7d69-40f3-a84d-66c4064b2833 · outbound
Exploration by Random Distribution Distillation Optimistic exploration even with apessimistic initialisation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e3eac3ba-11d2-43d6-b0c5-8b6136567d87 · outbound
Exploration by Random Distribution Distillation First return, then explore,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b9fd3f3-e076-46f5-9c9b-41d5b6359191 · outbound
Exploration by Random Distribution Distillation Maximum entropy gain exploration for long horizon multi-goal reinforcement learning,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5cf857af-89e4-4ca2-a879-eab0f299244d · outbound
Exploration by Random Distribution Distillation Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c604e84f-e73f-4779-abe6-859ab606b975 · outbound
Exploration by Random Distribution Distillation Vime: Variational information maximizing exploration,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3bd1767d-20b9-46ab-8076-546fff92419d · outbound
Exploration by Random Distribution Distillation Latent world models for intrinsically motivated exploration,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c153d218-9473-46f1-8e88-c661bd118573 · outbound
Exploration by Random Distribution Distillation Byol-explore: Exploration by bootstrapped prediction,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d6c8d62c-a2e3-4536-b407-aec3c56af37c · outbound
Exploration by Random Distribution Distillation Near-optimal reinforcement learning in polynomial time,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 34c19951-1d20-401e-b832-a17c25c8eddd · outbound
Exploration by Random Distribution Distillation R-max-a general polynomial time algorithm for near- optimal reinforcement learning,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2b4eccc8-4442-4767-884e-e766d5c53d5a · outbound
Exploration by Random Distribution Distillation Exploration in metric state spaces,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2323462e-f6f0-42c0-ada8-2e5f4c6b7953 · outbound
Exploration by Random Distribution Distillation Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1296819d-3453-4251-9daa-f33d920beaa8 · outbound
Exploration by Random Distribution Distillation Large-Scale Study of Curiosity-Driven Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19455772-98cc-4e3b-90ad-c96112cfe9d9 · outbound
Exploration by Random Distribution Distillation Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90cb151d-33f2-4fd5-822e-d814e87262a9 · outbound
Exploration by Random Distribution Distillation Reward shaping for reinforcement learning with an assistant reward agent,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c4a967f6-3bdb-4d7d-88ca-0a9ac1b6e017 · outbound
Exploration by Random Distribution Distillation Uncertainty-Aware Reward-Free Exploration with General Function Approximation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a784eea4-8f98-41fe-a5fa-92d32ed2579f · outbound
Exploration by Random Distribution Distillation Learning to shape rewards using a game of two partners,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 222d0b3e-5274-45e9-952d-26cd9b8ae3a4 · outbound
Exploration by Random Distribution Distillation Exploration-guided reward shaping for reinforce- ment learning under sparse rewards,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ad9875b9-beae-4235-9c03-555f8c462098 · outbound
Exploration by Random Distribution Distillation Addressing function approximation error in actor-critic methods,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b60a254a-37d9-46c1-8905-445957cd6893 · outbound
Exploration by Random Distribution Distillation Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8a5bb5-0f63-41ca-a63e-6709d318da05 · outbound
Exploration by Random Distribution Distillation The arcade learning environment: An evaluation platform for general agents,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 19200da9-cb16-4b58-9a2d-f90888c5e555 · outbound
Exploration by Random Distribution Distillation Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64d6c4b-226e-48cf-8329-d9d7ad39242c · outbound
Exploration by Random Distribution Distillation Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c2c5bf-4bf1-432e-a928-5136683d3d7f · outbound
Exploration by Random Distribution Distillation Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84cd675-cd39-47d1-84e9-e0182dc42bc3 · outbound
Exploration by Random Distribution Distillation Double check your state before trusting it: Confidence-aware bidirectional offline model-based imagination,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9fd9f4e2-1193-4ffe-8baf-8f0841a77004 · outbound
Exploration by Random Distribution Distillation Optimistic initialization for exploration in continuous control,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b2aa546e-f02f-4124-824a-6c104f9bbef4 · inbound
Exploration by Random Reward Perturbation Exploration by Random Distribution Distillation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.