Pith. sign in

Paper Citation Record · LEDGER

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

As of 21 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 2 inbound Pith citation observations for arXiv:2411.15370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15370 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:28:36.463613Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:29.076338Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T19:23:22.088073Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ae6fafe-f848-4081-b498-fde9a086bc11 · outbound

This paper cites an unresolved cited work.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:28:37.766178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.129633Z digest=sha256:96a0a8faa2d97b81952c6c99ca02c4f92bc368b932fd550cd7c403c1a6980aff

Observation fd06cfa7-acb9-42d5-95ec-84cbe33549f5 · outbound

This paper cites an unresolved cited work.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:28:37.755738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.178664Z digest=sha256:e32671265a5269d9f2eaf8c8fb90d9cdd678131964bfda69b0e9d6b6c8978ad1

Observation aa13e95c-d80e-479c-9b5b-07d3548afc0f · outbound

This paper cites If ˆY is an unbiased estimator of ¯Y and { ˆYj}j are i.i.d.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers If ˆY is an unbiased estimator of ¯Y and { ˆYj}j are i.i.d

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.664235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.258725Z digest=sha256:69a9c5bfb6549714308603732809ddef6ef8ed845b2faa3857213c6651f3d088

Observation 5ee2c12d-34a0-4015-8a43-4125c5648cc3 · outbound

This paper cites If ˆy is an unbiased estimator of ¯y and {yj}j are i.i.d.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers If ˆy is an unbiased estimator of ¯y and {yj}j are i.i.d

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.557407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.262384Z digest=sha256:42ba0824c535035be5da987893b23788f7ae15f89b50e4b5340545fda95e3af3

Observation 01d77d02-9cb7-4394-862f-b8e99688825f · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.546652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.266488Z digest=sha256:b125eb12de80d6312f26c09d36e7e3ee302decbf2ae7c51cd4dab022a2357e8f

Observation 4eb8622f-e1df-4424-95cc-a3f8043b31da · outbound

This paper cites Limitations.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Limitations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.424635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.270752Z digest=sha256:073738c9c2119ce9b8abf269ad62c06172f46c3a1d76086c324ad3a15ad74a19

Observation 4c169bb7-0ce9-404a-965f-bdaf10f68f15 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.317713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.274285Z digest=sha256:0e4413fe223958f5bd92a7248087044944af70aab778320ae96bef56b899e13f

Observation 43353a88-73fa-42b4-b7cd-100aff96d0fc · outbound

This paper cites We provide pseudo- code and implementation details which are easy to follow and reproduce.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers We provide pseudo- code and implementation details which are easy to follow and reproduce

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.307848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.277778Z digest=sha256:bd5282905bd567f89f5c1c6c24e2cc1e864f3c249dfb0aafaa0a7e80f8a4155b

Observation 9701a506-9e26-4dcc-846e-5aef48b1adc7 · outbound

This paper cites All data can be generated during training.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All data can be generated during training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.198589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.281382Z digest=sha256:f633bdaa7fb3fa144222b635cd3ef2583fcccd117da8ba59413fdfa7437056c1

Observation aabcf87c-1da2-4887-be5a-a45e22a20e85 · outbound

This paper cites We also list important hyper-parameters, neural network architectures, and other training details in the appendix.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers We also list important hyper-parameters, neural network architectures, and other training details in the appendix

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.045620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.285818Z digest=sha256:6dcd9ca4574f7318cf810daa3d152927d20a44a67028d6a60564b6c0613d44d2

Observation 7bb0f572-ce8a-49cc-8adf-b4389de3351f · outbound

This paper cites All results are averaged over 30 runs and reported with 95% confidence interval.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All results are averaged over 30 runs and reported with 95% confidence interval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.035305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.289957Z digest=sha256:dba1b2cfd046bc9a8aac74408d44c5ec277c54c7a731f8d0f367132729a70667

Observation b26a0333-19a7-444f-904e-9258f11ec5ff · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not include experiments

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.024001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.301253Z digest=sha256:3cddfbab393aec0e3b5b64ee955f153806bb5cd5a22a1b98e56b8954b0987166

Observation 7b98a8d1-a70b-4352-b53a-97f9a4d3dea4 · outbound

This paper cites • If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers • If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:37.013746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.391274Z digest=sha256:b47c4976f792a06cde1e381e82b1352dec47cb27f38246eac5de2532543acb87

Observation 598d7125-4740-43fe-868a-b06a68f200dd · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.759056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.446111Z digest=sha256:211edd413f02353fd77f3be5d249e86142ef2d3281383772bd54de65f8abc25a

Observation 62fa513f-08b8-4688-a0d1-7da0336bc486 · outbound

This paper cites an unresolved cited work.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:28:36.748818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.449703Z digest=sha256:91716735abca617e07dad1eb629c37e6e392be5ea50e1c83224f025f73182faa

Observation 3da4fab0-88f2-450e-931d-1eae4bd66df1 · outbound

This paper cites All our results are generated during training.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All our results are generated during training

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.738384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.453035Z digest=sha256:762a79f82fef54f698db74a75712c3d054950b8fc9ba062f334abd1d3fa44f23

Observation a2515985-b436-4cf4-91ef-c5c390052cce · outbound

This paper cites • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.728013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.456421Z digest=sha256:964ee1a62f3475760ad95cf25ee738a281efbdf548ac562ba02768497fa5d1c4

Observation c11a8980-e900-4a47-a2d6-18fdee074516 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.716860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.460059Z digest=sha256:0fbeaa710f2cad061068bf0d8af8be2c4deb32c7c59fd75cde7f137f0c6db1a7

Observation 095d6426-efb0-4455-8c3e-f0f3bbc49eba · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:28:36.503641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:28:36.463613Z digest=sha256:02eeef82b3811ac125a7db0c681bb303f69f2aeabf4a69cd6a26be813c28663a

Pith citing papers

Observation 5f805f2b-8114-45fc-9733-87f9dd8099b4 · inbound

The Hive Mind is a Single Reinforcement Learning Agent cites this paper.

The Hive Mind is a Single Reinforcement Learning Agent Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:23:22.090290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T19:19:27.381408Z digest=sha256:b5a4878795cbadd21d734de6b6b8b84a9efd02614b1250bee599b2e58c46f6eb

Observation cd65b16c-a140-415e-b947-da1f44889529 · inbound

Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning cites this paper.

Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:29.076338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:16:29.076338Z digest=sha256:4d3793c7e93b05f7121a074eb3ef124ec5e14de2eaebb353e0504cb48ecb8fce