Pith. sign in

Paper Citation Record · LEDGER

A Survey of Exploration Methods in Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2109.00157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.00157 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:58.574784Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 028e9a40-e4cf-43b0-87e4-ed52803edc9a · inbound

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy cites this paper.

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy A Survey of Exploration Methods in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:12:10.452960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:12:10.452960Z digest=sha256:3472ff36f92d005c0f2a433e80d9e616a1cdbef2a66bc57d78c987f82e8171e9

Observation 595fc39b-f931-4b4d-abbb-d9a20d1d694f · inbound

Effective Reward Specification in Deep Reinforcement Learning cites this paper.

Effective Reward Specification in Deep Reinforcement Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.307309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.307309Z digest=sha256:21f29fcaed84f45a939f450f55863e486153a28d7cd38da6d6a0655cfcdbc827

Observation 52181ded-a820-4ff4-a78a-693a6396b17b · inbound

Active Inference and Human--Computer Interaction cites this paper.

Active Inference and Human--Computer Interaction A Survey of Exploration Methods in Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:59:02.571787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:59:02.571787Z digest=sha256:ed931723a60011f4943e8f0e4135ade26b0fa8c81433337b600647a42f60474a

Observation 4c4361b5-ff42-4174-8b5a-f2b15909f195 · inbound

The impact of intrinsic rewards on exploration in Reinforcement Learning cites this paper.

The impact of intrinsic rewards on exploration in Reinforcement Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:11:48.173645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:11:48.173645Z digest=sha256:6fcf201bf3b0aef193d4adc08d1d9c24ea6a93da85f382b4e72f3d4b76a9181f

Observation 43a209fe-3b9d-41c8-957b-fbb07b00cbc4 · inbound

Parameter Estimation using Reinforcement Learning Causal Curiosity: Limits and Challenges cites this paper.

Parameter Estimation using Reinforcement Learning Causal Curiosity: Limits and Challenges A Survey of Exploration Methods in Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:58.574784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:58.574784Z digest=sha256:c12d7a08c1cfe0b8c69303a59637f2708ea4b8dd80612f637b6adb2c6cc21711

Observation 7e36c451-4f03-4bec-af1c-54f73c8859ec · inbound

Exploring Large Action Sets with Hyperspherical Embeddings using von Mises-Fisher Sampling cites this paper.

Exploring Large Action Sets with Hyperspherical Embeddings using von Mises-Fisher Sampling A Survey of Exploration Methods in Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:54.730532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:54.730532Z digest=sha256:662b548a37e54fd00c0dcd4e5a8a377cd71fac6a9c8927c28824e90305c78b33

Observation 3cd80d02-d104-44df-a2c3-ee39a4cb426c · inbound

Towards Human-level Dexterity via Robot Learning cites this paper.

Towards Human-level Dexterity via Robot Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:16.231424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:14:16.231424Z digest=sha256:8f88a8936400c9ed5bfa9b0c2fafe2f7985333e2284d79c69ebdc1fedf98a483

Observation 62f72c5a-6235-4654-bc26-fe632c7dce51 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey of Exploration Methods in Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.419668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.419668Z digest=sha256:2d42096232f81fb0776bb40e3c539691dd598d77ad199a5864b73f1ea7a2524f

Observation 7c8d4574-d36e-4a66-9513-d4dc796d50d7 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.597599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.597599Z digest=sha256:3453b774e2f9b8ef2b215acd78478ad5bf0257e91a7ae97caf289114a0991176

Observation 877220da-97a2-4565-9c7a-a2f10d61f3a2 · inbound

Smart Walkers in Discrete Space cites this paper.

Smart Walkers in Discrete Space A Survey of Exploration Methods in Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:47.682510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:19:47.050591Z digest=sha256:70eb669ef9f638cf08981d7c167efcf70638af2999b68f6eafa15a0c894f77fa

Observation 0ae1acbb-c158-48d9-bb24-f7539e748d95 · inbound

SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis cites this paper.

SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis A Survey of Exploration Methods in Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:50.844784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:06:10.718673Z digest=sha256:5f59f046dad8484695746effafe032008a8de2060460ffa63d388ac3f22b6efd

Observation 3413dc2e-46f0-4972-9d79-0f866fc922a9 · inbound

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring cites this paper.

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring A Survey of Exploration Methods in Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:15:31.460106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:13:16.406160Z digest=sha256:b14ccc9c9d0093ff43b99b1d060801e31ff70a6dda7b59a588deaae9298e49e2

Observation 2864721b-faa5-461b-a298-3e582951cbdf · inbound

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring cites this paper.

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring A Survey of Exploration Methods in Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T21:15:13.179473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:15:13.179473Z digest=sha256:110cf6892cf2b486fa1d84cb4a029b93195d0cb5a5c474a068601f6bc50c6457

Observation 6fbaf575-13ae-449d-af90-d33c5b0a82b7 · inbound

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling cites this paper.

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.278744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T09:39:29.875977Z digest=sha256:6ac022a49a359d8ce056f23e2beaff4ebc5b72e05134c594817e9d8f6d014168

Observation c62a7275-2b45-4015-8da9-98638dbab921 · inbound

Optimal Semiparametric Dynamic Pricing with Feature Diversity cites this paper.

Optimal Semiparametric Dynamic Pricing with Feature Diversity A Survey of Exploration Methods in Reinforcement Learning

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:06.944659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T17:34:49.811590Z digest=sha256:ab69b0c20df0bc8a87681f55f1c478cd36090129b62086c23c57100c0e7b4f88

Observation 768a09fc-41e6-4135-8ef3-2afd43b35674 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.653383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:f9734a0019542a959c81511ce3f72813e7802ac0bbaeaf5f541607886d8d39f9

Observation 5e0b398b-60e6-4157-aa44-89cf8ebba0e2 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback A Survey of Exploration Methods in Reinforcement Learning

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.089880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.089880Z digest=sha256:cceb8c6e846d27e48f33d82f8602eeeccf1e3e6e1405c90de043a537adfb41ff

Observation 463ddaed-3d44-4da6-9ce5-80f49d67d6bd · inbound

Parameter Exploration for RLVR via Variational Learning cites this paper.

Parameter Exploration for RLVR via Variational Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.730067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.730067Z digest=sha256:90cd4b490fcfa4cac1a68e3fd85c2feb05892ba0e1ddbff204adfd654efa3ea0