Pith. sign in

Paper Citation Record · LEDGER

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2501.12627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12627 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:31.896668Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T12:15:08.304150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.723680Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd12ae0c-9e7c-47c1-ba72-949a942c50d9 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.671849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.664680Z digest=sha256:ddbb833ce5ce49d8f73acdcfb5df26add9eb131199fa9c1790dbc3a4d2e7060c

Observation 777d5ba1-acf0-467c-9d2e-420e08bb6f61 · outbound

This paper cites Atari-5: Distilling the arcade learning environment down to five games.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Atari-5: Distilling the arcade learning environment down to five games

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.655050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.670067Z digest=sha256:abd8af919284c87a59172a12f1f236c6374afe044b20c20ed162f3d22a3a029e

Observation b025b990-4e9d-4136-91ad-7e30cf51b870 · outbound

This paper cites Existence, relatedness, and growth: Human needs in organizational settings.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Existence, relatedness, and growth: Human needs in organizational settings

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.639893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.674886Z digest=sha256:6901422b0afe53a4e89ba3f65e7d09a17ed3070451e4c28f8acd716ee068b755

Observation fab2e43a-96cf-48ae-900d-3b37447c7a19 · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Using confidence bounds for exploitation-exploration trade-offs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.624106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.679787Z digest=sha256:f887a6a92b7ce7efb49264484ae9b7b36bd60ddc5b1118100ccfa11cd9b3b914

Observation 42f52b22-413a-4239-ac7e-ae3ce297336c · outbound

This paper cites Never give up: Learning directed exploration strategies.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Never give up: Learning directed exploration strategies

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.608864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.684579Z digest=sha256:bb515d35a841c097b1df4969d7d52648e083dfab53a0e716c6949a80ab63d853

Observation dc46cc3b-895d-4b3d-8121-fd8c24dbfd99 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model The arcade learning environment: An evaluation platform for general agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.593250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.689560Z digest=sha256:7e5fcc578f6f0d2ed20ac5a76b70c74e5815071139bc82dcbef86799a984896d

Observation 1553f4ef-748c-4999-a2d7-b454bfa9b33e · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Unifying count-based exploration and intrinsic motivation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.578496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.694801Z digest=sha256:80ade0661207e42ea92a856a1902c4d3fe5756da536381610c7bb6fb2f7e2df2

Observation 29042ab4-e56a-47ea-a5dc-495eafe0d9fc · outbound

This paper cites A markovian decision process.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A markovian decision process

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.562582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.699397Z digest=sha256:c1833fa011ef6d6a79f7090b4dbc19e2c25e888668c7ae92dd5384adddc295ae

Observation 9d1afa23-e649-41e7-83cf-a20d2800927d · outbound

This paper cites Exploration by random network distillation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Exploration by random network distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.546712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.705179Z digest=sha256:5b04dbf358ec93608b2a234f0e7cfc8ba0bdf67709dd707e64f808a1cf3ab749

Observation 80fddff9-f795-4460-b351-96a5501b9a18 · outbound

This paper cites Explore, discover and learn: Unsupervised discovery of state-covering skills.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Explore, discover and learn: Unsupervised discovery of state-covering skills

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.531059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.710552Z digest=sha256:a4f7ef6d8762f3b3495bac47c492f72aefdf1d6933d4484aee48da6df3415bb5

Observation 10b4dfd5-e1dd-453f-95c2-e8e02292dbfc · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.514663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.715540Z digest=sha256:3d7eb4546b7968deb87495ad4720a3cd2ab55a35161dc649f8112fe08737e502

Observation 85ba021e-a615-4c9e-9059-04ffd982424a · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Leveraging procedural generation to benchmark reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.497491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.720712Z digest=sha256:77732994391d280b9ea13a4e31cd7acbfa3195ecc70421b5e4919da1d5328be0

Observation b48be9ad-2e89-446a-857f-a7f686b64864 · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Stochastic linear optimization under bandit feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.479715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.725530Z digest=sha256:588beae5d709115d0ba610983f4a88d317c009ff1e5d0eaa9c571a78f53cc74b

Observation 0dabe982-1812-4302-9b1a-383c4fc6b0af · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Diversity is all you need: Learning skills without a reward function

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.464273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.730698Z digest=sha256:f58113decbb7a296ba212bd17e5760eabc85c6e0e438b177669fe39e250fbe22

Observation d996408c-e6e1-41f5-9480-c491fba766b4 · outbound

This paper cites Adversarially guided actor-critic.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Adversarially guided actor-critic

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.448723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.735368Z digest=sha256:3349e284b81d7348e03b8b10c2bcd1f3566cfa321c84a2ff1f47d35ae64fe1ee

Observation c02a526d-a851-4081-a54a-850e1efcf554 · outbound

This paper cites Variational Intrinsic Control.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Variational Intrinsic Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.740034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.740034Z digest=sha256:a61ff865b7c4fc4088f0b369c76be73af3360dc06fb059e4ed89d14d8e45198f

Observation 02f8d2c1-69c4-40ef-b290-4a966b455c1f · outbound

This paper cites Fast task inference with variational intrinsic successor features.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Fast task inference with variational intrinsic successor features

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.432093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.745680Z digest=sha256:420bd6b05b9d80724f61f9ff5662ec9185ff6e5063469d7a2987005799eeaf81

Observation 40c39387-4a1a-4ff8-a700-7ce1765cfaae · outbound

This paper cites Provably efficient maximum entropy exploration.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Provably efficient maximum entropy exploration

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.414030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.750439Z digest=sha256:aa35ab3cf8e21fd3451f371d3fb5213aa6411a54e21b9f63715ce8066cf8384d

Observation 6c77159e-7fab-45ef-a465-6146559a7f82 · outbound

This paper cites Exploration via elliptical episodic bonuses.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Exploration via elliptical episodic bonuses

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.396886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.755324Z digest=sha256:0730e4f2d0676811387f9b31351d4c07ea5925cba4e1c93260511cd4e8d8e356

Observation ecdcb10c-28dd-4b67-8e33-85840a9f89e1 · outbound

This paper cites A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:04:32.011275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.760641Z digest=sha256:ca4e6a4c2e4459c818720654159360c1c232efb3756425eb7b7c224321adb94f

Observation 726615b0-541b-47b7-85c3-e284cc80958b · outbound

This paper cites Planning and acting in partially observable stochastic domains.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Planning and acting in partially observable stochastic domains

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.380601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.765980Z digest=sha256:f4ffe97d787e94a11c69dd4c23fade475d6d471b42350a019a025f0f242b982e

Observation bb2441da-0e0b-4c45-b9b7-92c3c81f1039 · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Curl: Contrastive unsupervised representations for reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.365128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.770736Z digest=sha256:5134921193afced131ee2cd4fa3f9f9c45d1ffa81baad01f7421f6f7837ae2fb

Observation 2b80cb0b-265f-46b1-a407-77add88a25d5 · outbound

This paper cites Cic: Contrastive intrinsic control for unsupervised skill discovery.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Cic: Contrastive intrinsic control for unsupervised skill discovery

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.349486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.775401Z digest=sha256:d1a58c8aa65d3eb0b5bf4f3162c93707fca66ef819858ff785ff5fc83cf897fe

Observation 0b57d581-455a-4b5d-9b79-286e760ea403 · outbound

This paper cites Urlb: Unsupervised reinforcement learning benchmark.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Urlb: Unsupervised reinforcement learning benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.333107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.780024Z digest=sha256:3b0f4caa5c1cf50420d69d8c9a67437b457c10580c703b69e56e11f0304fd9a3

Observation fbfafc83-6a9c-4275-aec1-51134eeed3cd · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A contextual-bandit approach to personalized news article recommendation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.785023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.785023Z digest=sha256:67d1c06232dae6742ed22c189d74e2c68fc0e2684ed9c12d4c373b4c85e8770c

Observation 586a0342-e2ce-4c97-a19a-f321767e38ac · outbound

This paper cites Aps: Active pretraining with successor features.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Aps: Active pretraining with successor features

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.305952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.789842Z digest=sha256:8152d07ccaa0f5eb4e2313c96807c170a51bfd47a3733a379eb3f37bcd6cbeb7

Observation f9c13c34-23d2-40fa-9e4e-a9a9e1d6d75e · outbound

This paper cites Count-based exploration with the successor representation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Count-based exploration with the successor representation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.290227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.794533Z digest=sha256:7db72cb702808cdf917ba69821f61695a36d396e7f6931e44fbb606aa3a5f6d5

Observation 167bf6db-1c33-49a3-a2db-506bafe739d2 · outbound

This paper cites A dynamic theory of human motivation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A dynamic theory of human motivation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.274824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.801075Z digest=sha256:aa2d5bc6ca74f129fec9c45d6e1aa038643e23b7d2b4d4c75095c7bf74f3ccb9

Observation 3e37be31-db05-4499-93aa-fec1ee4c345b · outbound

This paper cites Improving intrinsic exploration with language abstractions.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Improving intrinsic exploration with language abstractions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.258772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.806136Z digest=sha256:f7cb185550a46c341f78082a44c8a886edab0510e15428a59784a5047cf836d8

Observation f0e6ea44-e187-4119-b82d-106f8e2f4acb · outbound

This paper cites Count-based exploration with neural density models.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Count-based exploration with neural density models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.242416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.811005Z digest=sha256:e6719085be2f6319a89f4115c596fb02ddbb381047ba03ba45e5e94fb1d92837

Observation 486ab38a-5025-4edb-b7e2-680f42fbb893 · outbound

This paper cites Lipschitz-constrained Unsupervised Skill Discovery.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Lipschitz-constrained Unsupervised Skill Discovery

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.816155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.816155Z digest=sha256:4d20cd7f3b79fe909780b15e3c92b95b62eb6785c26e2978aca5b2d2468906a3

Observation d051dbc2-2728-4121-bad0-cf8043d7f72c · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Curiosity-driven exploration by self-supervised prediction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.820821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.820821Z digest=sha256:d84c1abe0a61598b3893d0cbab51af43115ba29f94e784480a891805d211410b

Observation dae2e2b5-728e-483f-a7c1-cdf48f88d5ec · outbound

This paper cites Self-supervised exploration via disagreement.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Self-supervised exploration via disagreement

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.216136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.825285Z digest=sha256:5d895252dc3a4af1ded8985faf069a7c94180d27708e2866ec0e740b2d5ae83c

Observation c253f65d-1884-4552-b4aa-34d7aeb1d127 · outbound

This paper cites Ride: Rewarding impact-driven exploration for procedurally-generated environments.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Ride: Rewarding impact-driven exploration for procedurally-generated environments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.199778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.829753Z digest=sha256:eca5e7f6c5a545f2c881433eb8aaa72d1e360bd8e431fb2707e75f16520da1fc

Observation e7f619b6-92d2-4171-836b-d697f8a232cb · outbound

This paper cites Minihack the planet: A sandbox for open-ended reinforcement learning research.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Minihack the planet: A sandbox for open-ended reinforcement learning research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.182309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.834609Z digest=sha256:218d6a3ec763ed5c5a78c9954b70abfc4323ef3d540c4e612e445de7f4c90140

Observation 51ccc0d7-1fd0-4e2b-b73b-846d65d96fa9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.839875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.839875Z digest=sha256:c2d32a683eb4c1cfc30caa7f79c29c05ac19b4f3d53a9be08d550fded6afb4d7

Observation e019307f-01e9-42f1-9754-8ce787f3b7cb · outbound

This paper cites State entropy maximization with random encoders for efficient exploration.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model State entropy maximization with random encoders for efficient exploration

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.166131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.845382Z digest=sha256:14315adbeced8533897c0a0ff866c0ed80f5d5082a19c71e852308e42ccce742

Observation 3bd73490-d8c9-4f91-baf4-0b9fe55ed8fc · outbound

This paper cites Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.850563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.850563Z digest=sha256:658553552f0c5feef7914409e5f67c78fdc516b829dddcdd5132946027afdb32

Observation d7c5f284-6475-4019-b3f9-cefc29542115 · outbound

This paper cites Reinforcement learning: An introduction.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Reinforcement learning: An introduction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.855932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.855932Z digest=sha256:3aeba3ab006fb99903561ca200291bd6828b014ac77718d73b7a8970877a25ee

Observation 2cef02a3-edf5-43c6-9fc8-4a27ada1fc55 · outbound

This paper cites \# exploration: A study of count-based exploration for deep reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model \# exploration: A study of count-based exploration for deep reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.860690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.860690Z digest=sha256:fed9c64b403bd1969da3086a38b04c35335f3bed9b02e8944aeaf79389a7c2e6

Observation b2bcbb78-6710-494f-8f9a-a84688d5851a · outbound

This paper cites Reinforcement learning with prototypical representations.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Reinforcement learning with prototypical representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.127782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.865517Z digest=sha256:63ef448ec6aa4894d5233f7915663b5f7205516439eb58028cd186a40b6ab7a3

Observation ce6a0bb6-ad98-4bdf-8008-b911a70e1984 · outbound

This paper cites Rewarding episodic visitation discrepancy for exploration in reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Rewarding episodic visitation discrepancy for exploration in reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.105311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.870168Z digest=sha256:61c041e7c7943fa4f404d3626c12e5aa3dbe72ac20bbbeba317a9980d7398c22

Observation 3f4e8c4b-f613-45e3-8860-13a13a60dd75 · outbound

This paper cites R \'e nyi state entropy maximization for exploration acceleration in reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model R \'e nyi state entropy maximization for exploration acceleration in reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.089219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.874687Z digest=sha256:de8ef06917b72ad228d789239f3f28f2b42c88c9cb8a668882f53535bc32f21e

Observation f3e1a87a-8f8a-406a-abc1-4777764238e2 · outbound

This paper cites RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.880479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.880479Z digest=sha256:5368d31143c55b99c54d362bcf99f8cf3b045009a1702097dd9cca2437435ea0

Observation 892dd3d4-1673-431a-b114-762c796b0303 · outbound

This paper cites Rllte: Long-term evolution project of reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Rllte: Long-term evolution project of reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.072848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.886926Z digest=sha256:f977e3fbeda1bfcf273b8082510fb3995bb704a22b2689382513248d1427cbe3

Observation c9f54259-2184-4fac-8134-bcbf9206621d · outbound

This paper cites Noveld: A simple yet effective exploration criterion.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Noveld: A simple yet effective exploration criterion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.057202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.891809Z digest=sha256:ef229c16dd26f252e5dc3e3152ceb11adf87c884353738f99d91e5520657e797

Observation a0e2ef72-6f7e-4268-8b52-866ea9ee5639 · outbound

This paper cites write newline.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.896668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.896668Z digest=sha256:b1ffb0bb54c11fd25cc0c37d648b3ae6503b16a4bedd961e70cdecdbfab64a49

Pith citing papers

Observation ea666f1f-983f-42cd-9966-94b2d78a1c1d · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Deep Reinforcement Learning with Hybrid Intrinsic Reward Model

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.724925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:423db877e6e21946f99824d449a7d259a88b657b1637e1b08d6435d7fdfe041d