Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:32:20.217328Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 3 inbound Pith citation observations for arXiv:2502.07523.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:32:20.217328Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:06.014406Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T12:46:27.860463Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75263cad-6299-44ef-8c73-abf37c6bf9d1 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Deep reinforcement learning at the edge of the statistical precipice
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ed9c9d59-5e49-495c-961c-2af5be4e7484 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization On Warm-Starting Neural Network Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f6ac76-8cfe-462c-bf34-87e6816846f2 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Layer Normalization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7b4c7d-f0f2-4df3-be89-633f0f8faffb · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization CrossQ: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d6659d78-ad05-4bb0-a86d-47dacf7dba7b · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Towards Deeper Deep Reinforcement Learning with Spectral Normalization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e046cc-1d8e-4d19-aa0e-5d1a8422f4c5 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef71f28d-42d5-40fc-97f9-c1beb070e77a · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e4f8f6-24f0-4a9f-a7ae-39ec91ea58a0 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Randomized ensembled double Q- learning: Learning fast without a model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 352a4d70-af65-4ce7-ada4-5c408307db7a · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Sample-efficient reinforcement learning by breaking the replay ratio barrier
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7388153-db91-4c99-bf03-b1b9f5fc9d3b · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Weight clipping for deep continual and reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 124445d9-ca17-48a4-ab62-79762922e9e2 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Jordan, Joseph E
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc1bb07f-f41d-4ef8-80ef-5542c77c39cf · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Sharpness-aware min- imization for efficiently improving generalization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93bc73ee-9cbb-49af-b36a-1a4d74c1b7c1 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Soft Actor-Critic Algorithms and Applications
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890de0fc-32be-4f12-9a3f-9a323e487a83 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Dream to control: Learning behaviors by latent imagination
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ab10e9bd-6f4a-4044-ab68-b771cc5d82d0 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d95cfb06-dee2-455e-815e-27f16afdc36a · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Learning continuous control policies by stochastic value gradients
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 775bbe21-ce93-4a3f-99e7-d3fe6ff23a9a · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Dropout q-functions for doubly efficient reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35af7e7d-9a2a-44d5-a273-7c282b4b0a45 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Normalization techniques in training dnns: Methodology, analysis and application
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a1b4d42-adff-45e9-8b04-bdba794b3256 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158c42e5-cd72-4b2b-83fa-4e16bf8ffe27 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1ea665-fb5e-450e-b8d6-b77cebc78778 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization When to trust your model: Model-based policy optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d7cbeaed-d115-4233-8cca-5a863f876ad7 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Deepmellow: Remov- ing the need for a target network in deep q-learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 582992de-d84e-4821-a9eb-6b892d79df62 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Adam: A Method for Stochastic Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0d8d24-004f-4f43-8235-6b05288739c5 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization JAXRL: Implementations of Reinforcement Learning algorithms in JAX, 2021
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0518ed5d-c1aa-4997-a5c8-434f8fe814ff · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 615629a0-8983-4261-a4d4-9821c9b73eb1 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Simba: Simplicity bias for scaling up parameters in deep reinforcement learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9b1a5874-9527-46ec-95fa-62312d3922b7 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Efficient deep reinforcement learning requires regulating overfitting
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 446266a6-c383-4f5f-abad-5374223ee8e8 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Decoupled Weight Decay Regularization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60866d2-054b-4aac-8e9d-4fe2dc0984d7 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Normalization and effective learning rates in reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9320d514-4261-4ba8-abd5-d14d6e858fb2 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Grokking deep reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db7f85ff-004f-48e8-a4ae-a052922d4b62 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a3bb533-3354-42f0-985f-4d7d98cb4fa0 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization The primacy bias in deep reinforcement learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b22ac48-c03a-4651-bfbc-f025972209f3 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e9d014-9436-43f3-a179-b5527363d280 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Markov decision processes: discrete stochastic dynamic programming
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79217d20-d803-4cfe-aea3-e7ed9c64138d · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Weight normalization: A simple reparameterization to accelerate training of deep neural networks.Advances in Neural Information Processing Systems (NeurIPS), 2016
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 38f87664-07ec-495d-b6ea-79eb253619ec · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Bigger, better, faster: Human-level atari with human-level efficiency,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd82cf46-10db-4510-a472-7e36d58917e2 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ddccf2aa-fed1-4b33-ada8-97c587bf4b0e · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e01e6f2-e993-479d-9e2c-a3a130f50738 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Sutton and Andrew G
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f15bc57-ea9e-49ba-816d-260ba7d40a00 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization DeepMind Control Suite
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229608f7-4f96-43a3-b34c-32edd1ff9e11 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Mujoco: A physics engine for model-based control
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0aa07f3-b903-4b0b-a8ee-26b53270196f · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization When to use parametric models in reinforcement learning? In Advances in Neural Information Processing Systems, 2019
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation caed7d6c-b0b2-48a8-8a35-6c0e6eaba9a9 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization L2 Regularization versus Batch and Weight Normalization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e034f50a-2aff-4ba1-806d-58df997a9bee · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c2e5cfb-aa5d-4209-abb3-c3fe1571c58a · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Root mean square layer normalization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9174a6a3-46b7-43ed-9071-aad4835e0467 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 714c4790-4185-46eb-9328-c49b541c7c85 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Limitations
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b3af780f-3a10-47f8-80fd-49a647ddfce4 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not include theoretical results
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4887074-6461-48b7-b3bb-c3d71997bd38 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization To aid reproducibility, we plan to release the code together with the camera-ready version of the paper
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b674ce1f-a0e4-425e-9a69-476838c88c3c · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization We plan to release the code together with the publication of the paper
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb87d9b0-aa13-4434-8428-bb10e92b1ccf · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not include experiments
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79634b99-c3fb-4dd3-bd36-d4f912570c3e · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Results are aggregated over multiple environments and 10 seeds each
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 81b417fa-0671-4dc4-aaa7-ce5973b1c6a6 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not include experiments
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b74d57b4-a312-4725-91d9-032675b7d9cb · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 73ed4cfe-9dba-41ea-8e2c-66046cbf7835 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization As actor-critic methods already enjoy a long history, there is no additional societal impact with this research contribution
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fbd4819c-cde6-425b-a14a-2103608adc12 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1cc4494b-c4ec-49df-94cb-8727302ce406 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not use existing assets
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1659f625-38a3-477a-8cb6-ec380efaf8b4 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not release new assets
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58009948-31b4-4bdf-81cf-8cdde5f88573 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05b1cf6-fef2-4e2e-807c-e4aa83e6e252 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization • Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7a401270-6cba-409f-870d-c645450b258a · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f2f1f8d1-f3ba-4ee3-a3e9-b0a4e6212f51 · outbound
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Bigger, Better, Faster: Human-level Atari with human-level efficiency
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47bb1096-30d7-43a5-a96a-02dbf597181b · inbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5df350-6c5b-40a5-8584-b9664a6ecbbb · inbound
Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3db05dd-9295-402a-bc36-ca5c22c7c0fe · inbound
Extending Differential Temporal Difference Methods for Episodic Problems Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.