Pith. sign in

Paper Citation Record · LEDGER

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning

As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2412.03767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03767 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:15:09.161752Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved37
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5334bbc9-0520-4a6c-ac8d-65fba9774625 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:06.702638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:06.702638Z digest=sha256:eadf676542f5ddeba8719b8feb46baca44a309520f14c3571da4375f353f363d

Observation 3f9ac66b-fcbb-4398-9916-83f6843b4059 · outbound

This paper cites Human-level control through deep reinforcement learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Human-level control through deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:06.751109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:06.751109Z digest=sha256:1cc662ea350b258f15e8fb513129aa998b6f4803e9cdd6f91590a65a026ed28d

Observation 2254d36b-be0f-4152-9df8-e6dc6447cd40 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Mastering the game of go with deep neural networks and tree search

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:06.774746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:06.774746Z digest=sha256:374a872f98d71137869ce23d8b4cee13387cd1cfb434bcfcff7b7a04cb5d01db

Observation e1e9b94e-931a-4d47-b4f8-88451eded4f8 · outbound

This paper cites Mastering the game of go without human knowledge.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Mastering the game of go without human knowledge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:06.834954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:06.834954Z digest=sha256:54ca0edee7a3e5c53fe7c4a43f7c5644638b9a8f9756ac023793e12685b8d337

Observation 628813d8-6d4a-4b89-9ea9-000295df3d8f · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:06.865821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:06.865821Z digest=sha256:8359bd905223133fcaae3b2522fc15e616f1ab8a98437962000ffc6988e128dc

Observation abd6207a-59ce-41b6-b65f-58e4c1187365 · outbound

This paper cites Alphastar: An evolutionary computation perspective.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Alphastar: An evolutionary computation perspective

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:15.094752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:06.902534Z digest=sha256:3d99ca5891073531a274cac3b5c59e2a912c331275ec601a71624c1367009bf3

Observation 8e6e462a-9ed2-46c4-b9da-9a4f92db2151 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unifying count-based exploration and intrinsic motivation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:06.974750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:06.974750Z digest=sha256:1d151ccae09e111d6926d0200e275c05877f14d2bc3215eaa13e4ca9893c5c30

Observation e641409a-cf0c-4a48-bc3b-ea8bf34a12af · outbound

This paper cites Curiosity-driven exploration by self- supervised prediction.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Curiosity-driven exploration by self- supervised prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.834946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.001141Z digest=sha256:1111d9336ed2482c5eb5ce4ef9e78b3d0aa5d72bef51bc1dd51aa8a71b4772f6

Observation 2329f066-004c-4d96-a2be-25cf83995d4e · outbound

This paper cites Count-based exploration with neural density models.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Count-based exploration with neural density models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.787945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.030124Z digest=sha256:7ca3c3cd2978760f0bff0f69cb2235c82cd599d9be39b7a0394fdccc7a64b629

Observation 64bacd47-a6d9-4b9a-9f86-fb6817e829a8 · outbound

This paper cites Exploration by Random Network Distillation.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.085548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.085548Z digest=sha256:67e2541f8b9bb6c51a791f74d57563298b4f7063431807e035b70c04fa9fb9f6

Observation 65300bd7-1be3-4533-a0fe-c8462a1fa556 · outbound

This paper cites Count-based exploration with the successor representation.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Count-based exploration with the successor representation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.175685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.175685Z digest=sha256:5ffd4e93b0c7650dbee8e6b818c11150fa3aa4a7234679b02f39f92a564e6bb9

Observation 68ef5268-d29d-4128-a877-3d8b8ac5d113 · outbound

This paper cites Self-supervised exploration via disagreement.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Self-supervised exploration via disagreement

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.574751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.234843Z digest=sha256:98c260a2ed13edf2c8db7862171b962177e01ec872381480ed5afda364c84118

Observation 47bcd4c3-323b-46c9-b3f5-4db9e47f67be · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Using confidence bounds for exploitation-exploration trade-offs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.294747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.294747Z digest=sha256:8e81e5ceb2358e78e66850b1e74435e86a91426f9af3c75472bf36f9250a0a91

Observation ab6e9ae4-2d5a-4412-9fc2-558f9265cf33 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Minimax regret bounds for reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.358712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.354760Z digest=sha256:485532d2088ab6e130df283853caee8391af8f09b45bbf6b95c7bc291fbe1873

Observation f7dfa20d-b28b-466d-812e-1002fb601f6c · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.426786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.426786Z digest=sha256:cbd40f55ed684c3d05c409ee68275efb5caa372b242cd64b511f0e10bc7f5599

Observation 04c70e77-3fdf-454a-a204-5f97e96e1c3e · outbound

This paper cites Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.164871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.464747Z digest=sha256:1d7495668617dc9c6a9037e781e1363ee1787043c83f35c538a63ee4b11f0e67

Observation 3b7e0732-f27d-4879-ba03-095be3993881 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Provably efficient reinforcement learning with linear function approximation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.087733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.506898Z digest=sha256:7df4721ffa31841bc090c03736e79c9501a66c144ba5694b4ea13d6200967fce

Observation 33bc6ba8-d99a-40b3-a9bf-90fad04cc456 · outbound

This paper cites Decoupling exploration and exploitation in reinforcement learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Decoupling exploration and exploitation in reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:14.004750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.560684Z digest=sha256:c966a681abce7c3c6501afce237946cedc4ba9d28bd25959da698f57af94a96c

Observation 8bd19351-5d4d-4e01-8c62-7f86af3a70d8 · outbound

This paper cites Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.599131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.599131Z digest=sha256:af8f32f9cd037e1807a77d7f19e87a2fd25bb11f7d443d1990ac505bd7677512

Observation d1f09285-4e35-4751-9c94-7a6d79a46543 · outbound

This paper cites A markovian decision process.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning A markovian decision process

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:13.884744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.624746Z digest=sha256:a38567b790474b0f20276764012dc6599ee989517b615b2ef98bf4704e7234ec

Observation 3f2ce05c-dd91-4d4b-9c40-10a93e2dbee2 · outbound

This paper cites Q-learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Q-learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.654754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.654754Z digest=sha256:5d3ddf3d15c4349137afc70297848024c2b4847ca2155975b52405f5b32e0283

Observation 128d976f-aa71-4d07-a54b-8d6d74e4b6c6 · outbound

This paper cites Sample-optimal parametric q-learning using linearly additive features.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Sample-optimal parametric q-learning using linearly additive features

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:13.693847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.702779Z digest=sha256:4930aa346e6ffea9880e971b442e3e851f5ee1a8e579c6c6d343dd400a938b5d

Observation 98021d54-31fd-4116-b296-44297a5eab4f · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.754751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.754751Z digest=sha256:b6f312795554cdd1bf0ae39610e832fe5eed373e4bd54225a0823be1a14684d6

Observation 85190290-751b-4b7f-bd7b-b1710f629b3e · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.794754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.794754Z digest=sha256:dd268b594d7b5362e06cc77d4f7ae892f0074cad83a725d32c5aea923c24212b

Observation 4e59b5a9-1f9c-46fc-a131-2ce1a955dc52 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.815965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.815965Z digest=sha256:ef20db65e96a4229424813c2ae25194fa78b346a3ef1faf0018b8c13f02af222

Observation 05a83d6f-1fd5-4a22-b5ca-f2c75d4819bf · outbound

This paper cites On the likelihood that one unknown probability exceeds another in view of the evidence of two samples.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.856198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.856198Z digest=sha256:620e6c3f639dfbae4a1ad0047e69a6da290ed832fd0f416cd7d7de68080d7836

Observation 4b104fa7-23b0-41f8-a669-27613a683d74 · outbound

This paper cites Intrinsic motivation systems for autonomous mental development.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Intrinsic motivation systems for autonomous mental development

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:13.410220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.890468Z digest=sha256:975e5823cba4cb49522cf2e5eb7d51a58e8ca263d674e1214310cfaf39508c71

Observation faf025e1-c39c-4958-876f-773b523613b0 · outbound

This paper cites # exploration: A study of count-based exploration for deep reinforcement learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning # exploration: A study of count-based exploration for deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:13.374746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:07.924750Z digest=sha256:1aba866cb01b7545052fa81704f7388bdd5a4980e09a7e2f9b6861d0b27d797f

Observation db2679fb-1c8f-4e62-9d24-1f5c9217bc7b · outbound

This paper cites DORA The Explorer: Directed Outreaching Reinforcement Action-Selection.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning DORA The Explorer: Directed Outreaching Reinforcement Action-Selection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.964749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.964749Z digest=sha256:d44fb3720dcfd81d0b03cbcf288dc9aa8a4096d0eb20dab73b1093ff7ca26f24

Observation 0002ce96-1da6-4f36-84cb-ec4a8d0951fc · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Go-Explore: a New Approach for Hard-Exploration Problems

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:07.996689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:07.996689Z digest=sha256:c14e34fe52189115438f6f6319892544bc6a05ca29f1a2e0921a123e32584384

Observation 1a2de18c-5b05-43eb-a53b-d391695f7c9d · outbound

This paper cites Bayesian reinforcement learning: A survey.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Bayesian reinforcement learning: A survey

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:13.264224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.011642Z digest=sha256:bb1635c2ccd0ceea3e2c1df58cead05475dd4225f1944c73b76b28ca0fd6f726

Observation 60c46138-98ad-4aa7-b6d3-44fde8f74f58 · outbound

This paper cites Noisy Networks for Exploration.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Noisy Networks for Exploration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:08.036093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:08.036093Z digest=sha256:ea93c30dacc374f8d337a630408dbc81b006949f94be67230604b6122535e058

Observation bd4eb3cb-e7ce-48cd-a173-7862acf5787b · outbound

This paper cites Deep exploration via bootstrapped dqn.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Deep exploration via bootstrapped dqn

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:08.066238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:08.066238Z digest=sha256:e23a15a066f60dc3ec8d36b34598cabe256f8f6a0ff4047b62b186c850b3211f

Observation e7ab1784-c3be-4c0c-955f-70b573cb2ec2 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:08.104763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:08.104763Z digest=sha256:f413b8bc23c344573db4bfa5eda1a9660a7cad319a855d84884c68dc6da5983d

Observation bbfcc985-d3fc-420c-b456-82a1cf45a88c · outbound

This paper cites The option-critic architecture.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning The option-critic architecture

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:08.144746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:08.144746Z digest=sha256:da5f76ce60729327a020a2dd936976eb6ae3b5b4b310d2565b14361748e0562e

Observation f4aaee55-e2ac-4197-9a3c-a0591c7617a7 · outbound

This paper cites Temporally-Extended {\epsilon}-Greedy Exploration.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Temporally-Extended {\epsilon}-Greedy Exploration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:08.194750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:08.194750Z digest=sha256:55cc8c6c8802e1f6a4367fc7d9ad9b42de360a6ed0fe9212a60f178386146fd3

Observation 271d509a-fdfa-44ae-80c1-671978738de5 · outbound

This paper cites Redeeming intrinsic rewards via constrained optimization.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Redeeming intrinsic rewards via constrained optimization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:12.934749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.264752Z digest=sha256:c4c014764a287dca1a57ef678a530157b830032ef182f165ab2a2e0973432e51

Observation 71e915c1-a226-46d4-aed6-19f27206c9ca · outbound

This paper cites LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:15:08.324750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:15:08.324750Z digest=sha256:e3768f4aa8dc9ae37901a2c80dccc4fe202f4cac123d0a1a94e0776018dce981

Observation 4712356e-d74f-4682-abf0-fac63198ccaf · outbound

This paper cites Decoupling exploration and exploitation for meta- reinforcement learning without sacrifices.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Decoupling exploration and exploitation for meta- reinforcement learning without sacrifices

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:12.800760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.375104Z digest=sha256:2f5d0726b82c5ecd047836e9d0c4e8f310ddf7757ebd6aaa845d0b2165810dcf

Observation b508ad01-c083-4f05-9667-f25dc12f9852 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Improved algorithms for linear stochastic bandits

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:12.681661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.424747Z digest=sha256:1aff4b94c2a36071948f6fddaf653c6f84d2b46c79fadd182f5c5cbf49b1c0d3

Observation 3b90f39c-5e0e-4d96-b646-1ad6e1f70fb1 · outbound

This paper cites healthy reward.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning healthy reward

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:12.544744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.474749Z digest=sha256:535a3a3da84b144004e0709ca14b65390a308f52ecf2b60dd233c9e0b08b8d79

Observation 64bbd336-8514-4a9e-9b96-3a84893009f4 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:12.394744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.534749Z digest=sha256:e5699afcf6c0a906e7ce22a44de481299fcb71ab64ddf335812b083f225a38af

Observation bb213bf0-c021-46e3-b007-99e2d45802cc · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:12.273545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.584750Z digest=sha256:0336037ee5a3000d83990b86745a5142a85666f5bdd14cb857ccf5ca874cfd83

Observation bbcfd74b-7ed8-4f46-a22b-763fb7180bd3 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:12.168873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.634752Z digest=sha256:9af4f89c539fb1883f587c3fcca7730e6847ed65521b6e3f59ff1043cfbbe901

Observation bb7bb341-8ed3-4940-8281-747c3d8bf1c3 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:12.044746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.684749Z digest=sha256:009bddfc91fbee98f7cc63a6e7a0f4b7722b429c58647804cf18149c5c95da7c

Observation 044fc1a1-a401-4154-8631-ac68a34efe74 · outbound

This paper cites And further, X k∈I ˆV k 1 (sk.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning And further, X k∈I ˆV k 1 (sk

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:11.934771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.715584Z digest=sha256:5b95d3ab8151e90f8ba1894aac8a57510a09388044bdb8bbb972f738f4db48cd

Observation 18b5cdc6-d7b9-4d7c-adf4-867574b3376d · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 47

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T22:15:11.854927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.754750Z digest=sha256:fc42abb8ddf0517ec3332bc65ce0207660b83a921ca79d0d4d94a5ca446eab89

Observation b0b6a364-8c63-42a2-9be8-af6df57fd8fd · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.764744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.783299Z digest=sha256:73b68c1dc43cf80bf11ecd84bfd5d527fcb637af8062bde38881f248092c656f

Observation f177c735-10fd-493b-8169-6ea3309ad979 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.686497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.814756Z digest=sha256:d5e37c98d121837bae02ae32e5b150c117e8d39507e5f8a79a881c627d0ee71e

Observation 27395cf3-b54a-47f4-ab34-03d0d1522bdf · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.644824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.854751Z digest=sha256:d5047b05ebcd286f15e49a0f9a23488001e24bee096ecaace8f3953075fe8b49

Observation 46a2843a-2a71-48c6-88e8-34561d8910d0 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.495901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.887250Z digest=sha256:500ff5c187bc81869c353a96f4e86c2a2d74346d2aa1fbba5cc1bd45b967df1d

Observation 9fbde005-7587-4409-9913-0737366ab294 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.385416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.937334Z digest=sha256:cd54a2ac618c67cd72f0ecf2f9e68e168bf8b05379f263b63001a298507b42c0

Observation bd9f997b-0b94-4f54-8333-d9aaf86b273f · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.295153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.947509Z digest=sha256:f4e58d5007f009439076480b0f0e568044b0f698cfe33eae523872f76de2c812

Observation ad613507-f506-49f9-a00e-6c2fec908dc8 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:11.124748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:08.974750Z digest=sha256:13470850c9d6e56469dc441edbcfc3fd67410082fa067d3d35246a686ff4b60a

Observation 6e636333-bf5f-418a-9df4-2b0bae31d1bb · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:10.963835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:09.004749Z digest=sha256:0fecc382c44b0733cac51bb49d4bbae42b9b7e030cc341933c4ceb9ba5832e7b

Observation 56747326-f28c-44ad-b205-df13f1e36f05 · outbound

This paper cites By recursively applying Equation.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning By recursively applying Equation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:10.834756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:09.044748Z digest=sha256:9bdd9f0258711661ba1ccbeb1389499456a0f3ce842e96fc3a5bfc51cbe37e94

Observation 1c08dec8-5c4c-4d53-b80d-a485250dfe38 · outbound

This paper cites an unresolved cited work.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:15:10.695308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:09.094752Z digest=sha256:36fcdcbb6523f5e909e121874d0be4969e3225d88727aeb6490b2a2a3655329d

Observation 4c669c78-15b3-4b06-9ff0-bb53113a9b73 · outbound

This paper cites Hence, the total regret is given by: KX k=1 ˆV k 1 (sk.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning Hence, the total regret is given by: KX k=1 ˆV k 1 (sk

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:10.560933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:09.119779Z digest=sha256:233defb58c763510948215fab6e1379621261185f83986f0a248fb8d331448e7

Observation 21551000-93d4-4bbc-881b-c4214708596e · outbound

This paper cites And equivalently, we conclude that our algorithm obtains ϵ-optimal policy with ˜O( d3H 4 ϵ2 ) samples with probability at least 1 − δ.

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning And equivalently, we conclude that our algorithm obtains ϵ-optimal policy with ˜O( d3H 4 ϵ2 ) samples with probability at least 1 − δ

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:15:10.425577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:15:09.161752Z digest=sha256:29562ac58fdfd64c9e0a3e04c9b8f2a3e7abf8359cc27ab66542938f4c972125

Pith citing papers

No inbound Pith citation observations are available.