Pith. sign in

Paper Citation Record · LEDGER

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

As of 19 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 2 inbound Pith citation observations for arXiv:2411.15370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15370 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:28:36.463613Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:29.076338Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T19:23:22.088073Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ae6fafe-f848-4081-b498-fde9a086bc11 · outbound

This paper cites an unresolved cited work.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:28:37.766178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.129633Z digest=sha256:d0ac9e1a9439bbe06f9e7211237cefa6dd1a982af89c73dfcf8983f730fa03e4

Observation fd06cfa7-acb9-42d5-95ec-84cbe33549f5 · outbound

This paper cites an unresolved cited work.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:28:37.755738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.178664Z digest=sha256:b2bd71f332e15eaabcc3520a29ba5211e09f4a54fc53b1801c49fc09d5d1c052

Observation aa13e95c-d80e-479c-9b5b-07d3548afc0f · outbound

This paper cites If ˆY is an unbiased estimator of ¯Y and { ˆYj}j are i.i.d.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers If ˆY is an unbiased estimator of ¯Y and { ˆYj}j are i.i.d

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.664235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.258725Z digest=sha256:416ea62d45278bbab1c3397477c3a2b49224861d89603ec0008a864a4b5dcf74

Observation 5ee2c12d-34a0-4015-8a43-4125c5648cc3 · outbound

This paper cites If ˆy is an unbiased estimator of ¯y and {yj}j are i.i.d.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers If ˆy is an unbiased estimator of ¯y and {yj}j are i.i.d

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.557407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.262384Z digest=sha256:3a57b3cf556cb5896e130c540e152638a1e388b2c42e07b06fbd009e540ed1bf

Observation 01d77d02-9cb7-4394-862f-b8e99688825f · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.546652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.266488Z digest=sha256:cdfbd2ef74af27938dc8a1992126f921ae38357e41de6a8d95dcf35be8861e22

Observation 4eb8622f-e1df-4424-95cc-a3f8043b31da · outbound

This paper cites Limitations.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Limitations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.424635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.270752Z digest=sha256:06a771942b7d873e2b20bc78b1a4231881d76bc6eb0af6f79a6cc9b1a2bc5451

Observation 4c169bb7-0ce9-404a-965f-bdaf10f68f15 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.317713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.274285Z digest=sha256:f3304fbd8333630be26ce3f3d9a4c46f1f1cfe6ba0d717ce13936e72fbcd6992

Observation 43353a88-73fa-42b4-b7cd-100aff96d0fc · outbound

This paper cites We provide pseudo- code and implementation details which are easy to follow and reproduce.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers We provide pseudo- code and implementation details which are easy to follow and reproduce

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.307848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.277778Z digest=sha256:b61ec3391de6bb34c0291a3568e66f8a9d24704066d02f94116951d0cdd056f5

Observation 9701a506-9e26-4dcc-846e-5aef48b1adc7 · outbound

This paper cites All data can be generated during training.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All data can be generated during training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.198589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.281382Z digest=sha256:ad916fd3534af5e747c99d7880553ac6876b77942b381c29d6ffbbc612239bd0

Observation aabcf87c-1da2-4887-be5a-a45e22a20e85 · outbound

This paper cites We also list important hyper-parameters, neural network architectures, and other training details in the appendix.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers We also list important hyper-parameters, neural network architectures, and other training details in the appendix

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.045620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.285818Z digest=sha256:d994df9da3e218116878faabd9b9e1886ac2f7c39a906e9f76165ff69771e82d

Observation 7bb0f572-ce8a-49cc-8adf-b4389de3351f · outbound

This paper cites All results are averaged over 30 runs and reported with 95% confidence interval.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All results are averaged over 30 runs and reported with 95% confidence interval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.035305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.289957Z digest=sha256:69822c690cda9e36b344774a64e0014f26a347e8b718cfdc14fc70f7725e01cc

Observation b26a0333-19a7-444f-904e-9258f11ec5ff · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not include experiments

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.024001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.301253Z digest=sha256:f972a66e4dc4c6581dced973bfad6667fa91c5604b8535b755f677a73d2f646a

Observation 7b98a8d1-a70b-4352-b53a-97f9a4d3dea4 · outbound

This paper cites • If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers • If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.013746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.391274Z digest=sha256:ef4996a8da1b8922a0138b606bebea36c6ee5f1e4422af0ba8b5e64222620f4b

Observation 598d7125-4740-43fe-868a-b06a68f200dd · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.759056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.446111Z digest=sha256:e952e9e452d5b6a8d97839f60e67b186adde8b7d7b0205514a3b74d392969e2d

Observation 62fa513f-08b8-4688-a0d1-7da0336bc486 · outbound

This paper cites an unresolved cited work.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:28:36.748818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.449703Z digest=sha256:e092f43753ab351af403e03ed23df2b1da4b0461707bb539c90772009853b811

Observation 3da4fab0-88f2-450e-931d-1eae4bd66df1 · outbound

This paper cites All our results are generated during training.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All our results are generated during training

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.738384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.453035Z digest=sha256:bfd2c2c09454fb783f3905391fdd2077eff17fc84c29cbb464ab059209ce6629

Observation a2515985-b436-4cf4-91ef-c5c390052cce · outbound

This paper cites • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.728013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.456421Z digest=sha256:1b3cf31f787224978f3f2a543152e4de0a12efa691b6a70cdefb7948486cd1f5

Observation c11a8980-e900-4a47-a2d6-18fdee074516 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.716860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.460059Z digest=sha256:b507099acd987c7dbcc78045eb7258f57aac91ecb4d490ec40ffb195352d93dc

Observation 095d6426-efb0-4455-8c3e-f0f3bbc49eba · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.503641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:28:36.463613Z digest=sha256:007f962a86872c22b72b823235853bd12a5969730248e9f4ebe5dbfe377d945a

Pith citing papers

Observation 5f805f2b-8114-45fc-9733-87f9dd8099b4 · inbound

The Hive Mind is a Single Reinforcement Learning Agent cites this paper.

The Hive Mind is a Single Reinforcement Learning Agent Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:23:22.090290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T19:19:27.381408Z digest=sha256:2e3763773512ec2a706008d152d1d4b6808241f68dd352ef8fadbaaefa117fa9

Observation cd65b16c-a140-415e-b947-da1f44889529 · inbound

Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning cites this paper.

Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:29.076338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:16:29.076338Z digest=sha256:a04806426941dd95dc83ba05f98b1770894a9488f314d27f767cdfee74896894