Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.06937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06937 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:56:51.152652Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be150886-de80-4762-a6a8-42644b6bad6c · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.596693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.042565Z digest=sha256:709048e7bdcf2a4d1d9aabb3461b3dfe6ad84c25d09fd8af42909387cb043f20

Observation 93c12656-9030-4f74-8e1f-8cba2be7ec58 · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks G., Naddaf, Y., Veness, J., and Bowling, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.587835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.046887Z digest=sha256:0eaeaa146ef3badd4d542bffa16b3f61db05f1e645fdf10daf7c157465b73770

Observation d14898ed-0568-4795-a409-d931006dbf8c · outbound

This paper cites OpenAI Gym.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.050601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.050601Z digest=sha256:cf3467dc5e7322d9bdbe10ee439611279c577d3b701437869f445e182ad8136e

Observation cbe57015-1b5f-430c-b418-b5b2bd07c762 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.579136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.055134Z digest=sha256:65726bab5aa64828e1c41e50b2d78f423f4ca8ae771a9d7659860b237f372ad4

Observation d2fb5b3c-811c-4c25-bd61-bc90e63e0b63 · outbound

This paper cites Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.059025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.059025Z digest=sha256:c06c93a29336655cf055e7a6c45beeecd5633b0e45be2ffa34a8f7c9f44614af

Observation 4da91d0c-c520-49ec-9ea2-c098833d9a8e · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.062911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.062911Z digest=sha256:a4023702622dc208c7b36ee49aa44d6e0493814e2fc0d3208cf82c17d8b0ef65

Observation e803cdbb-79e9-457f-8208-ec2546712ee9 · outbound

This paper cites and Petrik, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks and Petrik, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.564149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.067089Z digest=sha256:aa5d486775d60b2382fd36e54dabd19cb36a368decab86f3d5b1b158f1ff2121

Observation cbebf486-942b-40e4-9eb7-2f4cadf725a9 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.070550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.070550Z digest=sha256:e38234a3dfc545f572c00dc4a52b1640da7078b7b86f05c0d826697c1e8826bb

Observation 3577ed12-96a2-4ede-afe8-fbe2a3f4b295 · outbound

This paper cites RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.074552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.074552Z digest=sha256:462a4cd07431dbb51ffc3eb1980961c2cd12483ad9c191be935626b5ec3c94fb

Observation 5c5ea7ba-f326-4ded-8cca-75798e2af797 · outbound

This paper cites Continuous control with deep reinforcement learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Continuous control with deep reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.078679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.078679Z digest=sha256:a9c678f49f872ef9f7355026368689e7c80dd27a5354f36aa25be0484cbf2464

Observation 0bc6b9d6-5972-481e-acdd-4019b2ad8d9a · outbound

This paper cites Average-Reward Reinforcement Learning with Trust Region Methods.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Average-Reward Reinforcement Learning with Trust Region Methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.082564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.082564Z digest=sha256:2c4e02eebc6deb585a73d51fd902ac7642fb66fb44ac862f0e6d5ffd5c9bed02

Observation 835823a5-df91-487b-a6f5-ab970128cb0b · outbound

This paper cites A., Veness, J., Bellemare, M.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks A., Veness, J., Bellemare, M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.087603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.087603Z digest=sha256:f4c346f2f47cde13fd1cd97f62dda527e76cb3e3d559df1fc54b84852c653baf

Observation 9f9ef73d-43e4-433f-be92-dde05275b5d1 · outbound

This paper cites Reward Centering.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Reward Centering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.090784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.090784Z digest=sha256:1429ebdadd31a43bf687e7c53b76f9dbf663072ded4cfa0fa585526d34a995c6

Observation 9822b3bd-4d1c-4ac9-84f5-6723d589abaa · outbound

This paper cites Jelly Bean World: A Testbed for Never-Ending Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Jelly Bean World: A Testbed for Never-Ending Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.094911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.094911Z digest=sha256:ca79cd7e665b40e21ff631701977b8d5ec229cf62440558818db1a00f3b2e2a8

Observation b3a91d80-5572-4fc1-a42c-e842cb5e85ba · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.549244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.098506Z digest=sha256:16e3e85380f76d6574183ffbb2ed506e0dbfae2ad31fe6031a6495d3f1c2c24b

Observation b351cda6-3d10-4448-9656-c9523f29e2ad · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.538881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.101851Z digest=sha256:d0d967a0fbc82d82eb9a98048b5be8487ba288cfce40908bd77e890c66e450ae

Observation 519c400b-9380-42fe-89b6-375b81f9ba7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.104986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.104986Z digest=sha256:961d7841d4a51763d0600b24510824dcf5461707098c02f873c263835f65858b

Observation f765d528-5c76-4c55-892f-ea08522312a1 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.528853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.108332Z digest=sha256:c411049caf524ff97b40e3ff45b483c9996b896d3439fe16311253ae677dc665

Observation 7d38b37d-f827-444a-9ddb-b1a33c365145 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.518642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.111420Z digest=sha256:9d778785aac264d2955c5b39a76f42b2761986efddd8fb6c685209fd61a296c1

Observation fb175955-ecca-48b7-8b67-7ef70bbd0fc0 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.114743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.114743Z digest=sha256:c9f6a6f71825b5d91c167d6e460d96dd3b24192afc49f367dcb738f6598f8e1b

Observation a58ef605-33d2-43b5-97bb-51781700d1d0 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.117972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.117972Z digest=sha256:800086dcb7a0644ed5be4201099f515c47fa0d70e5e6d9cfc8e511f008a31e87

Observation 3fcd1410-718a-4941-88da-70582473c74f · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.121085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.121085Z digest=sha256:c865860bcb27ea738ed8f2d1fc2e720f34c0e40f90824d029c7b4972310b4306

Observation 32099d3d-34fd-4afb-9ad3-424c62f33400 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.498595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.124304Z digest=sha256:2fdaf6266ad78235d6f78ab89e4a324aa47a6835db97760e6f82199a4203a444

Observation c768aa35-d9c8-4e8c-b19c-23112d861c4f · outbound

This paper cites and Ross, K.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks and Ross, K

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:56:51.488242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.127557Z digest=sha256:72d4aaa0aa2adf4855ab58be6c261bfcb29a7438f72322224ec037defc6a7c51

Observation 05517a7e-0d4f-4860-88f3-f56349d0ec11 · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:56:51.477254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.130679Z digest=sha256:bf4b8a515ae8262fb072728fcd74429172b67168ccd628e2f85003537cb17507

Observation 3c0fb13f-2a7f-4cd3-ae6c-34f461c50e48 · outbound

This paper cites The Ingredients of Real-World Robotic Reinforcement Learning.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks The Ingredients of Real-World Robotic Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.133757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.133757Z digest=sha256:e33049725604b73a9100a9be2f1b469bdbfd49ddf5d8ed574f57094a32db46ca

Observation 47f1ddd2-023b-4c39-8c16-9f6ae5e4777a · outbound

This paper cites Pearl: A Production-ready Reinforcement Learning Agent.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Pearl: A Production-ready Reinforcement Learning Agent

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:56:51.352949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.137298Z digest=sha256:1cc40ee6c77afef2a05ef56956f28cc60bc1ff55a579996ce2ad8a9d7232bc59

Observation e88cee22-f087-413d-a8ff-26293f0d462b · outbound

This paper cites write newline.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.141171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.141171Z digest=sha256:575cc555f6ca968ecf4c5674832b331ff2ae3138702b7d7d0eac6ff2576cce9f

Observation 432d02ac-23d8-4d74-b0e4-b4f2566d4be1 · outbound

This paper cites @esa (Ref.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks @esa (Ref

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.145296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.145296Z digest=sha256:3e40fabcff91d05bf407bacd370f636f8c74038429bbf4d05dcc5e4dc4a5904f

Observation 0c86f820-0284-474a-ba05-541e0913f5ac · outbound

This paper cites an unresolved cited work.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:51.149172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:56:51.149172Z digest=sha256:dd6b613dd79f3a910aee841462eb683d30c4ac72dddec505c3703f4f152581b3

Observation 152e618d-2879-4b74-bba0-39abe64e30a6 · outbound

This paper cites yF࡞g5. t]k e_kx kϣ], [#>3,>Mj <3oseӡ؎O޲ 7v 10 ΓZ Snc ay xط<Xت֬2/̡ Z̄峖G?Y[x=S c _ZSFX3#v)7n֎a | t 6IͰ|q֎Ğb /eev .Jj 6 4^ D[OVY L |axj ?#+#ڱ 44 . T c 7 W O &# O` !ң ծ [bY ncMj .0^?.

An Empirical Study of Deep Reinforcement Learning in Continuing Tasks yF࡞g5. t]k e_kx kϣ], [#>3,>Mj <3oseӡ؎O޲ 7v 10 ΓZ Snc ay xط<Xت֬2/̡ Z̄峖G?Y[x=S c _ZSFX3#v)7n֎a | t 6IͰ|q֎Ğb /eev .Jj 6 4^ D[OVY L |axj ?#+#ڱ 44 . T c 7 W O &# O` !ң ծ [bY ncMj .0^?

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:56:51.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:56:51.152652Z digest=sha256:e93ea846abb8a1f73cc653290742d19146188d754a0dfeaad4d76ebc4b504361

Pith citing papers

No inbound Pith citation observations are available.