Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:28:36.463613Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 2 inbound Pith citation observations for arXiv:2411.15370.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:28:36.463613Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:29.076338Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T19:23:22.088073Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9ae6fafe-f848-4081-b498-fde9a086bc11 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd06cfa7-acb9-42d5-95ec-84cbe33549f5 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aa13e95c-d80e-479c-9b5b-07d3548afc0f · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers If ˆY is an unbiased estimator of ¯Y and { ˆYj}j are i.i.d
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ee2c12d-34a0-4015-8a43-4125c5648cc3 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers If ˆy is an unbiased estimator of ¯y and {yj}j are i.i.d
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 01d77d02-9cb7-4394-862f-b8e99688825f · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4eb8622f-e1df-4424-95cc-a3f8043b31da · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Limitations
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c169bb7-0ce9-404a-965f-bdaf10f68f15 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not include theoretical results
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 43353a88-73fa-42b4-b7cd-100aff96d0fc · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers We provide pseudo- code and implementation details which are easy to follow and reproduce
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9701a506-9e26-4dcc-846e-5aef48b1adc7 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All data can be generated during training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aabcf87c-1da2-4887-be5a-a45e22a20e85 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers We also list important hyper-parameters, neural network architectures, and other training details in the appendix
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7bb0f572-ce8a-49cc-8adf-b4389de3351f · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All results are averaged over 30 runs and reported with 95% confidence interval
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b26a0333-19a7-444f-904e-9258f11ec5ff · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not include experiments
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b98a8d1-a70b-4352-b53a-97f9a4d3dea4 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers • If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 598d7125-4740-43fe-868a-b06a68f200dd · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that there is no societal impact of the work performed
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62fa513f-08b8-4688-a0d1-7da0336bc486 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3da4fab0-88f2-450e-931d-1eae4bd66df1 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers All our results are generated during training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2515985-b436-4cf4-91ef-c5c390052cce · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c11a8980-e900-4a47-a2d6-18fdee074516 · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 095d6426-efb0-4455-8c3e-f0f3bbc49eba · outbound
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f805f2b-8114-45fc-9733-87f9dd8099b4 · inbound
The Hive Mind is a Single Reinforcement Learning Agent Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cd65b16c-a140-415e-b947-da1f44889529 · inbound
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.