Pith. sign in

Paper Citation Record · LEDGER

Exploration by Random Distribution Distillation

As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.11044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11044 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:10:12.221295Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:59.951972Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:08:00.455034Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0eb65fc1-9d28-4857-9db5-a9974d1e5a56 · outbound

This paper cites Deep exploration via bootstrapped dqn,.

Exploration by Random Distribution Distillation Deep exploration via bootstrapped dqn,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.994412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:11.977476Z digest=sha256:af6cd166eb0f0fd2a83278e1f90058808deb508b58621d9b668c02fee24eba32

Observation f016da28-af73-4286-b8e6-a27bb6062496 · outbound

This paper cites An analysis of model-based interval estimation for markov decision processes,.

Exploration by Random Distribution Distillation An analysis of model-based interval estimation for markov decision processes,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:11.983113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:11.983113Z digest=sha256:c9af8475b0a65ec2b31eb02d0038fa5cfafe14f2209b201c467f17d11bf14800

Observation a61e4d41-e894-48fd-b0ab-a672e6ba11ad · outbound

This paper cites Minimax regret bounds for reinforcement learning,.

Exploration by Random Distribution Distillation Minimax regret bounds for reinforcement learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.966612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:11.989632Z digest=sha256:5fea94a066e3425b100814c80fb74096133e5ace046523d7786591cab844d8fc

Observation 9725abb0-e729-4bfd-8f64-7eeb29a7b53b · outbound

This paper cites The Alberta Plan for AI Research.

Exploration by Random Distribution Distillation The Alberta Plan for AI Research

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:11.995579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:11.995579Z digest=sha256:4e8d17303fbbe65965b4b86b9aef66e1e39363f2d32ff395836df72a61944c1a

Observation 1d3b377c-e0b8-4b17-9c1a-0d476d9926ce · outbound

This paper cites Unifying count-based exploration and intrinsic motivation,.

Exploration by Random Distribution Distillation Unifying count-based exploration and intrinsic motivation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.950687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.000973Z digest=sha256:d4159c15d45748f2c125630f196445f5b72624ef006bef2860a52e00eae789a6

Observation 817e8a24-ba9f-49fe-8186-7c05d023cdf9 · outbound

This paper cites Flipping coins to estimate pseudocounts for exploration in reinforcement learning,.

Exploration by Random Distribution Distillation Flipping coins to estimate pseudocounts for exploration in reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.932671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.005610Z digest=sha256:c139af44a9f8c2aab4c86c7b78dfddc3aeb6fcbdc1dae5d30e9c8c2d1710f31b

Observation 516956b7-2f51-493f-8ad2-2d9ce29e0446 · outbound

This paper cites Count-based exploration with neural density models,.

Exploration by Random Distribution Distillation Count-based exploration with neural density models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.912587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.010837Z digest=sha256:27e40d1e11c8dcb442b107359d5b73d19f5e55838ca2012cd2b834d78e6c2d8a

Observation 8eb97b88-0ec9-4cf7-8bfb-4439e6cb9798 · outbound

This paper cites Count-based exploration with the successor representation,.

Exploration by Random Distribution Distillation Count-based exploration with the successor representation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.895576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.015671Z digest=sha256:7d042ae8ce2bf1c255b630aafc801a4b0fb7edcbaa4b0ef0fce61951a1329d93

Observation 8ca4efab-9682-45c3-98d6-2f9852e4cf09 · outbound

This paper cites Curiosity-driven exploration by self- supervised prediction,.

Exploration by Random Distribution Distillation Curiosity-driven exploration by self- supervised prediction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.875598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.020460Z digest=sha256:6a2d15ab3ee01e46ced73defe86626f4a1606e6c8ab4778a87e606e4d05f3b2d

Observation a6877a50-f0f6-43a4-9779-cc36f7fa25b5 · outbound

This paper cites Exploration by Random Network Distillation.

Exploration by Random Distribution Distillation Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.025295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.025295Z digest=sha256:bb044c4c569145138ec9e8c89738fb1c32a3b49dd38eed210bea1f6d79e95f1a

Observation 0aaf07ee-afad-4a18-bc13-95d8f3f07399 · outbound

This paper cites Is q-learning provably efficient?.

Exploration by Random Distribution Distillation Is q-learning provably efficient?

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.856842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.030592Z digest=sha256:64054f796720c9328d24a55686d32f1ab339b552a31a4e291669de152a944ad5

Observation 8115b221-7f28-4551-9109-4bb4ae93f11b · outbound

This paper cites Exploration and anti-exploration with distributional random network distillation,.

Exploration by Random Distribution Distillation Exploration and anti-exploration with distributional random network distillation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.838173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.035218Z digest=sha256:e604ecaa675775143c0df2cdac8296fa79cae56b4fb2baaf5e1aad5a68daa6b1

Observation 6615c679-dfe5-454b-a919-25d1a598fbb3 · outbound

This paper cites Q-learning,.

Exploration by Random Distribution Distillation Q-learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.039603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.039603Z digest=sha256:6c2206e46c72f12342f91ddc590d54177adf401a052cf935aaf054be3a5e4ec7

Observation fcc62367-2383-4049-8f92-ef7e0c2e6340 · outbound

This paper cites Human-level control through deep rein- forcement learning,.

Exploration by Random Distribution Distillation Human-level control through deep rein- forcement learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.810716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.044184Z digest=sha256:c6307d66d71c176e9a58d11d384a3f3328e18d602d20ab723a20562e05330e1f

Observation 94a3152a-ea97-4944-85ae-e239488fc964 · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning,.

Exploration by Random Distribution Distillation Rainbow: Combining improvements in deep reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.793387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.048799Z digest=sha256:45b22a4685b102e4722df849dd4b76cc3a0d638222593a933dc8b9fa0265ba28

Observation c4edd3f9-b9a4-4f3f-b4b9-0ac936755f1d · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,.

Exploration by Random Distribution Distillation Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.775280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.053695Z digest=sha256:4b1bc94e79d2445a6b3841d5a2ee5b9ea1ddafdd8e7de40b081f7558e0ded7f5

Observation fd5aa7e3-bdec-4dca-b5e0-599adafbf90c · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs,.

Exploration by Random Distribution Distillation Using confidence bounds for exploitation-exploration trade-offs,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.759499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.059293Z digest=sha256:8a26175ba53bd9c84576428c2d5ec470ad07ec4a1189729db046c5718954c8a1

Observation 71522f23-58ac-4361-9bb3-4d77e6c40594 · outbound

This paper cites Improving generalization for temporal difference learning: The successor representa- tion,.

Exploration by Random Distribution Distillation Improving generalization for temporal difference learning: The successor representa- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.743270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.064338Z digest=sha256:ec9ca82343808d1f03f7a3db0812c57fca8c331243796a3ab2f8978a2fa607d7

Observation 97705774-3838-4cb9-b896-fc0e311ec1b2 · outbound

This paper cites On Bonus-Based Exploration Methods in the Arcade Learning Environment.

Exploration by Random Distribution Distillation On Bonus-Based Exploration Methods in the Arcade Learning Environment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.070120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.070120Z digest=sha256:a93f878358a9292093e06e426e5adc49d56d2aa304ea6f658255cc307028c854

Observation 5734298d-bb8e-4b35-b5f2-1a368f586323 · outbound

This paper cites # exploration: A study of count-based exploration for deep reinforcement learning,.

Exploration by Random Distribution Distillation # exploration: A study of count-based exploration for deep reinforcement learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.726622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.075831Z digest=sha256:b90f9cfc7ad45d2f52f08261ecafaedb5c3087871281d9e0fa9cb37d77aaed52

Observation 59e9d073-7d69-40f3-a84d-66c4064b2833 · outbound

This paper cites Optimistic exploration even with apessimistic initialisation.

Exploration by Random Distribution Distillation Optimistic exploration even with apessimistic initialisation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.710904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.080291Z digest=sha256:100ec334dc2a50209a489e936a79cd3604498926bb528aeb7831746935187db8

Observation e3eac3ba-11d2-43d6-b0c5-8b6136567d87 · outbound

This paper cites First return, then explore,.

Exploration by Random Distribution Distillation First return, then explore,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.084814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.084814Z digest=sha256:6b99738b295af3f0290d7df7264a071eae3a79223d24b0c4aeef2c48e7e3de59

Observation 7b9fd3f3-e076-46f5-9c9b-41d5b6359191 · outbound

This paper cites Maximum entropy gain exploration for long horizon multi-goal reinforcement learning,.

Exploration by Random Distribution Distillation Maximum entropy gain exploration for long horizon multi-goal reinforcement learning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.684813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.089055Z digest=sha256:365ed2a5193a72df09b2806a534ac0869ad1c94d727173407e7e5674067fe8d8

Observation 5cf857af-89e4-4ca2-a879-eab0f299244d · outbound

This paper cites Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models.

Exploration by Random Distribution Distillation Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.093833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.093833Z digest=sha256:dc0eb3e3af6bc5f93e8ff4758c2a5291a847cf9a759a6fae825d489bcb2cd907

Observation c604e84f-e73f-4779-abe6-859ab606b975 · outbound

This paper cites Vime: Variational information maximizing exploration,.

Exploration by Random Distribution Distillation Vime: Variational information maximizing exploration,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.667413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.098404Z digest=sha256:8bbd7bca5bb1d0327b5123de501d28eb02c3867203e98b06ab737273da62170c

Observation 3bd1767d-20b9-46ab-8076-546fff92419d · outbound

This paper cites Latent world models for intrinsically motivated exploration,.

Exploration by Random Distribution Distillation Latent world models for intrinsically motivated exploration,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.647994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.103056Z digest=sha256:8163387372f76b07b4a3d7df116b566309cd9f67803cf169f633b466bde8a629

Observation c153d218-9473-46f1-8e88-c661bd118573 · outbound

This paper cites Byol-explore: Exploration by bootstrapped prediction,.

Exploration by Random Distribution Distillation Byol-explore: Exploration by bootstrapped prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.629848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.107954Z digest=sha256:b72f1094f326c7f8a54b855d0872a4cbf4e47f3bfdb260a9b02b1d76a6f992f4

Observation d6c8d62c-a2e3-4536-b407-aec3c56af37c · outbound

This paper cites Near-optimal reinforcement learning in polynomial time,.

Exploration by Random Distribution Distillation Near-optimal reinforcement learning in polynomial time,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.613760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.112695Z digest=sha256:cab87b23cf8ac954d9800cf6bc736e7792727b9e3a12248ca90ef8516edfb0d2

Observation 34c19951-1d20-401e-b832-a17c25c8eddd · outbound

This paper cites R-max-a general polynomial time algorithm for near- optimal reinforcement learning,.

Exploration by Random Distribution Distillation R-max-a general polynomial time algorithm for near- optimal reinforcement learning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.597184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.117429Z digest=sha256:b0ecbc0908d8c9f5157dd8d0e59ece7170c6a60cda2582913f0013bdd58848ef

Observation 2b4eccc8-4442-4767-884e-e766d5c53d5a · outbound

This paper cites Exploration in metric state spaces,.

Exploration by Random Distribution Distillation Exploration in metric state spaces,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.582226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.123483Z digest=sha256:28180e2997e4531253bcb5d8e8925b3d2514989b4979b4c5e795c61bff3cd9cc

Observation 2323462e-f6f0-42c0-ada8-2e5f4c6b7953 · outbound

This paper cites an unresolved cited work.

Exploration by Random Distribution Distillation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.129378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.129378Z digest=sha256:6dd7a0388a9a004017fc5b3e383d40702d71317f0a458384063a156632432cf1

Observation 1296819d-3453-4251-9daa-f33d920beaa8 · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Exploration by Random Distribution Distillation Large-Scale Study of Curiosity-Driven Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.135139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.135139Z digest=sha256:1bec5c221c7e19a40fd3b59315eb27341161e9e3e72830297d0c300881139f2e

Observation 19455772-98cc-4e3b-90ad-c96112cfe9d9 · outbound

This paper cites Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning.

Exploration by Random Distribution Distillation Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.140204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.140204Z digest=sha256:269c220cbefdc6661541b393feee23db94123454545fc81a4cd5f8b8c864aa7e

Observation 90cb151d-33f2-4fd5-822e-d814e87262a9 · outbound

This paper cites Reward shaping for reinforcement learning with an assistant reward agent,.

Exploration by Random Distribution Distillation Reward shaping for reinforcement learning with an assistant reward agent,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.555544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.145618Z digest=sha256:970fff7c1b6ec25fb1cb733716dfb37fb26e8c7ccd96b919eeb139c8de8960b2

Observation c4a967f6-3bdb-4d7d-88ca-0a9ac1b6e017 · outbound

This paper cites Uncertainty-Aware Reward-Free Exploration with General Function Approximation.

Exploration by Random Distribution Distillation Uncertainty-Aware Reward-Free Exploration with General Function Approximation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:10:12.345576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.151636Z digest=sha256:bc25b2fc0a95bee10c709a72a2d7550b08eb4dd146bbaf115f741d6f0814bbe7

Observation a784eea4-8f98-41fe-a5fa-92d32ed2579f · outbound

This paper cites Learning to shape rewards using a game of two partners,.

Exploration by Random Distribution Distillation Learning to shape rewards using a game of two partners,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.539899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.158130Z digest=sha256:70574c970e6a85748554ba320b03707ec48b0dbf0ad5af11f00d750285c68cd7

Observation 222d0b3e-5274-45e9-952d-26cd9b8ae3a4 · outbound

This paper cites Exploration-guided reward shaping for reinforce- ment learning under sparse rewards,.

Exploration by Random Distribution Distillation Exploration-guided reward shaping for reinforce- ment learning under sparse rewards,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.523289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.164057Z digest=sha256:d3f91f007b54ef7e2fbfd0a4a5f3dc92d2c62b1f3d21baf2d3735536ed5734a4

Observation ad9875b9-beae-4235-9c03-555f8c462098 · outbound

This paper cites Addressing function approximation error in actor-critic methods,.

Exploration by Random Distribution Distillation Addressing function approximation error in actor-critic methods,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.508233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.168867Z digest=sha256:f6c0f41e45b68c4e2b856b63239241f43e498c455e0e17776984fda8278d4bca

Observation b60a254a-37d9-46c1-8905-445957cd6893 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exploration by Random Distribution Distillation Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.176055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.176055Z digest=sha256:be9c683284a5fdcb0bae30c085bf558da7cbc98574024820fc6cc18f51153ce9

Observation aa8a5bb5-0f63-41ca-a63e-6709d318da05 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents,.

Exploration by Random Distribution Distillation The arcade learning environment: An evaluation platform for general agents,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.493051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.180842Z digest=sha256:df915a9024c7e0ca7b73d0f871dae08c3504a8211366d86a440c16a4a9b12d60

Observation 19200da9-cb16-4b58-9a2d-f90888c5e555 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Exploration by Random Distribution Distillation Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.192143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.192143Z digest=sha256:86c1cc656e4fc4e96db9c60fc6cb0bb4ba99713f4e0e4e5ff3abdf7be99a000a

Observation a64d6c4b-226e-48cf-8329-d9d7ad39242c · outbound

This paper cites Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research.

Exploration by Random Distribution Distillation Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.197928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.197928Z digest=sha256:e280ea52e1a21258904b5b9a3cad2f049846e45668ec3593a62b44571f32e175

Observation a8c2c5bf-4bf1-432e-a928-5136683d3d7f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Exploration by Random Distribution Distillation Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:12.204381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:12.204381Z digest=sha256:9807d4860da272b622e164fa274489d39974ea6c5a2b08eb45705ac6d1dfd6c6

Observation b84cd675-cd39-47d1-84e9-e0182dc42bc3 · outbound

This paper cites Double check your state before trusting it: Confidence-aware bidirectional offline model-based imagination,.

Exploration by Random Distribution Distillation Double check your state before trusting it: Confidence-aware bidirectional offline model-based imagination,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:10:12.477164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.214743Z digest=sha256:cc2b77320eee5caf44803b2eb4b40bf8237f68cfc17c060e5a269a7aae6f738f

Observation 9fd9f4e2-1193-4ffe-8baf-8f0841a77004 · outbound

This paper cites Optimistic initialization for exploration in continuous control,.

Exploration by Random Distribution Distillation Optimistic initialization for exploration in continuous control,

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:10:12.461580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:10:12.221295Z digest=sha256:e61d581fd67f0584fe98a52507017f613689a4aaeb75f87dfb4d25d7d14c5545

Pith citing papers

Observation b2aa546e-f02f-4124-824a-6c104f9bbef4 · inbound

Exploration by Random Reward Perturbation cites this paper.

Exploration by Random Reward Perturbation Exploration by Random Distribution Distillation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:08:00.459145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:07:59.951972Z digest=sha256:7c21590d535295923a209a10d4489bc81fa8d98f5e039fb6fd426eda8ecb57f8