Pith. sign in

Paper Citation Record · LEDGER

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization

As of 9 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 3 inbound Pith citation observations for arXiv:2502.07523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07523 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:32:20.217328Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:06.014406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:46:27.860463Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact3
  • verified fuzzy39
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75263cad-6299-44ef-8c73-abf37c6bf9d1 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.108258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.945704Z digest=sha256:cc5559877bd6526c2d8162f6a59de98521a464e9ffcc8d4ac1e6338500b7812e

Observation ed9c9d59-5e49-495c-961c-2af5be4e7484 · outbound

This paper cites On Warm-Starting Neural Network Training.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization On Warm-Starting Neural Network Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:19.951110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:19.951110Z digest=sha256:2618a8e1702d5553506f555c992c6f92947daf1546fd7ad06656e11cda0c3a4c

Observation c9f6ac76-8cfe-462c-bf34-87e6816846f2 · outbound

This paper cites Layer Normalization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Layer Normalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:19.956080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:19.956080Z digest=sha256:0eeab58d557d65ef26bcf36efc436c9497ec48bdb641da0a59cd2407a08cce35

Observation 0c7b4c7d-f0f2-4df3-be89-633f0f8faffb · outbound

This paper cites CrossQ: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization CrossQ: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.094998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.960996Z digest=sha256:39c60367efed226953f28c55915a0f62740c110290284f7099c012870a16f840

Observation d6659d78-ad05-4bb0-a86d-47dacf7dba7b · outbound

This paper cites Towards Deeper Deep Reinforcement Learning with Spectral Normalization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Towards Deeper Deep Reinforcement Learning with Spectral Normalization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:19.966074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:19.966074Z digest=sha256:e328072da6d63a3fad526b0e60b5db3f64181a9e251d5516d11aca64bc32cb18

Observation 18e046cc-1d8e-4d19-aa0e-5d1a8422f4c5 · outbound

This paper cites Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:19.970952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:19.970952Z digest=sha256:d9bc8827df67d6fdcde77f39a00109ed0fbde97016abbb5b602d3442c95addcf

Observation ef71f28d-42d5-40fc-97f9-c1beb070e77a · outbound

This paper cites MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:19.976294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:19.976294Z digest=sha256:1041d64ee49ce108d27e1482ed28af1ed0723354e95a33642a48121dabee5500

Observation b7e4f8f6-24f0-4a9f-a7ae-39ec91ea58a0 · outbound

This paper cites Randomized ensembled double Q- learning: Learning fast without a model.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Randomized ensembled double Q- learning: Learning fast without a model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.081159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.980914Z digest=sha256:4fc58eba9e507cafd5d619e9a6f72c2e6a128eb0fe877129e9fdae0f7f1448b3

Observation 352a4d70-af65-4ce7-ada4-5c408307db7a · outbound

This paper cites Sample-efficient reinforcement learning by breaking the replay ratio barrier.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Sample-efficient reinforcement learning by breaking the replay ratio barrier

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.067453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.985373Z digest=sha256:fae4c1121ddb7befa128cca6fcbbb55d90e73285d404c1b83059628ea3008d0c

Observation b7388153-db91-4c99-bf03-b1b9f5fc9d3b · outbound

This paper cites Weight clipping for deep continual and reinforcement learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Weight clipping for deep continual and reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.054193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.989742Z digest=sha256:880497f51641ab2338c346595288426a5bb75885367445d84049a0b0b597e376

Observation 124445d9-ca17-48a4-ab62-79762922e9e2 · outbound

This paper cites Jordan, Joseph E.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Jordan, Joseph E

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.040767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.994044Z digest=sha256:3c54e3eb904a5fe1b504bc99783ae809f9de5c8fc59cc9447cca8e4b114d7401

Observation bc1bb07f-f41d-4ef8-80ef-5542c77c39cf · outbound

This paper cites Sharpness-aware min- imization for efficiently improving generalization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Sharpness-aware min- imization for efficiently improving generalization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.027624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:19.998499Z digest=sha256:9a7a5e57c16072d186a6598ae93bad94457f50a995090a62a9a083331791e69d

Observation 93bc73ee-9cbb-49af-b36a-1a4d74c1b7c1 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.003004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.003004Z digest=sha256:d346bd66bd796ea14912741367e25612114e3e3ccbe3a18ba280362578808996

Observation 890de0fc-32be-4f12-9a3f-9a323e487a83 · outbound

This paper cites Dream to control: Learning behaviors by latent imagination.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Dream to control: Learning behaviors by latent imagination

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.014128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.007777Z digest=sha256:bacf5d2b846dca11ccd62590920741c7c5e5f934144d53f6b959add3d58ae968

Observation ab10e9bd-6f4a-4044-ab68-b771cc5d82d0 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.012049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.012049Z digest=sha256:efb61198b8c0b949d9b41c4410547710672673f7760ce2b5815eca26a216cfb5

Observation d95cfb06-dee2-455e-815e-27f16afdc36a · outbound

This paper cites Learning continuous control policies by stochastic value gradients.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Learning continuous control policies by stochastic value gradients

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:21.000354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.016563Z digest=sha256:0ea59d2360eb512ccd53239bf6b8786e25e967037499cda77f8ead8e3af83867

Observation 775bbe21-ce93-4a3f-99e7-d3fe6ff23a9a · outbound

This paper cites Dropout q-functions for doubly efficient reinforcement learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Dropout q-functions for doubly efficient reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.986639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.020759Z digest=sha256:c6fcd57c014fcea50927716070dd1a6397b27ac56fd69387e934bffdb68f94ea

Observation 35af7e7d-9a2a-44d5-a273-7c282b4b0a45 · outbound

This paper cites Normalization techniques in training dnns: Methodology, analysis and application.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Normalization techniques in training dnns: Methodology, analysis and application

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.972649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.024954Z digest=sha256:79a3aef1fc14b081c4744fdbc38c0bc2cc6ce4b09dec1f5e43351b4d9d9793da

Observation 9a1b4d42-adff-45e9-8b04-bdba794b3256 · outbound

This paper cites Dissecting Deep RL with High Update Ratios: Combatting Value Divergence.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Dissecting Deep RL with High Update Ratios: Combatting Value Divergence

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.029295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.029295Z digest=sha256:f9906ef0f1784058598b2e8cf05bb100f0784ff666b1d764f360b476cb37328d

Observation 158c42e5-cd72-4b2b-83fa-4e16bf8ffe27 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.034194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.034194Z digest=sha256:4844a531d4339f42817b81b221dad4f12df75e708f045e062ba3c05d3b7a65b5

Observation 7a1ea665-fb5e-450e-b8d6-b77cebc78778 · outbound

This paper cites When to trust your model: Model-based policy optimization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization When to trust your model: Model-based policy optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.958602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.038685Z digest=sha256:1b81ddf0346f01101f408f30c94431829655c0f3ad365926f91671459ae96ddd

Observation d7cbeaed-d115-4233-8cca-5a863f876ad7 · outbound

This paper cites Deepmellow: Remov- ing the need for a target network in deep q-learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Deepmellow: Remov- ing the need for a target network in deep q-learning

Reference 22

Resolution
verified exact
doi, observed 2026-08-08T12:32:20.254991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.042944Z digest=sha256:1e48473e2b69e8c48cdc376e735ffcfd79c08a0012f16f397a485e8a435cf808

Observation 582992de-d84e-4821-a9eb-6b892d79df62 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.047225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.047225Z digest=sha256:4baf5c9ac61f7d11ee1bf1ae27988c3a709e7ac02c5edd1c81ef53b6b210aff2

Observation ef0d8d24-004f-4f43-8235-6b05288739c5 · outbound

This paper cites JAXRL: Implementations of Reinforcement Learning algorithms in JAX, 2021.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization JAXRL: Implementations of Reinforcement Learning algorithms in JAX, 2021

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.944655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.051755Z digest=sha256:790c7883ea7b8c203e6832679e486d44a2d2bbb9d3355a3bde153bc3c6eab8b3

Observation 0518ed5d-c1aa-4997-a5c8-434f8fe814ff · outbound

This paper cites PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:32:20.379411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.055893Z digest=sha256:85057a22e2d5d7552a0411b52800009b8f3bd9b0348bf167c52e63a01e885617

Observation 615629a0-8983-4261-a4d4-9821c9b73eb1 · outbound

This paper cites Simba: Simplicity bias for scaling up parameters in deep reinforcement learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Simba: Simplicity bias for scaling up parameters in deep reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.929383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.060235Z digest=sha256:e9f1ca27ea2a21dbeadc76bf72456c86c68ce4d57097dd11509e4fbcd81e21c8

Observation 9b1a5874-9527-46ec-95fa-62312d3922b7 · outbound

This paper cites Efficient deep reinforcement learning requires regulating overfitting.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Efficient deep reinforcement learning requires regulating overfitting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.913730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.064502Z digest=sha256:07f8f9883351893757e95b1f5b5b308ffbd7af26dd40c624b3557b5f178d55e2

Observation 446266a6-c383-4f5f-abad-5374223ee8e8 · outbound

This paper cites Decoupled Weight Decay Regularization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.068947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.068947Z digest=sha256:9be27fe169e34866b80acd41a4e2531bed10d95bd3882db7c51068c77cdb5cb3

Observation a60866d2-054b-4aac-8e9d-4fe2dc0984d7 · outbound

This paper cites Normalization and effective learning rates in reinforcement learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Normalization and effective learning rates in reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.899444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.073266Z digest=sha256:d7d81af2eb5a067d51b79febdf6f19dd3f259cba65cbcfc0c8e6d23bf33f717d

Observation 9320d514-4261-4ba8-abd5-d14d6e858fb2 · outbound

This paper cites Grokking deep reinforcement learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Grokking deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.885023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.077446Z digest=sha256:5e55e227a2cec44f3b97b23daf20f6a11f7499ad9825085d03f4fa3866c33a68

Observation db7f85ff-004f-48e8-a4ae-a052922d4b62 · outbound

This paper cites Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.871019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.081605Z digest=sha256:a8c08c73b36e8b91b5710869b8ca855bd69953d704843725ccda7b3f6b6b5637

Observation 6a3bb533-3354-42f0-985f-4d7d98cb4fa0 · outbound

This paper cites The primacy bias in deep reinforcement learning.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization The primacy bias in deep reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.856594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.085717Z digest=sha256:d8fb671cb30a979eb0685d9594b40aacf51bfb72227205a087cb63ed66caa946

Observation 1b22ac48-c03a-4651-bfbc-f025972209f3 · outbound

This paper cites Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.089815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.089815Z digest=sha256:14e7418e1fc423a3131c70e84a61cdf8c812768ee43b48265f68b5acb9243718

Observation 05e9d014-9436-43f3-a179-b5527363d280 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Markov decision processes: discrete stochastic dynamic programming

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.094222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.094222Z digest=sha256:1f30cd5278c229afe2b70446b8786c08a3c740862f050bba81d4bcd8606ee21c

Observation 79217d20-d803-4cfe-aea3-e7ed9c64138d · outbound

This paper cites Weight normalization: A simple reparameterization to accelerate training of deep neural networks.Advances in Neural Information Processing Systems (NeurIPS), 2016.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Weight normalization: A simple reparameterization to accelerate training of deep neural networks.Advances in Neural Information Processing Systems (NeurIPS), 2016

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.833567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.098434Z digest=sha256:94a27fe02991a02ea5d722ecc44e20d598a16d387859f1f6414a42762a5a3257

Observation 38f87664-07ec-495d-b6ea-79eb253619ec · outbound

This paper cites Bigger, better, faster: Human-level atari with human-level efficiency,.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Bigger, better, faster: Human-level atari with human-level efficiency,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.819182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.102964Z digest=sha256:be755fed83024dce0b573629af592e4625d5ba2640e28afab01d879719d9df06

Observation bd82cf46-10db-4510-a472-7e36d58917e2 · outbound

This paper cites Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:32:20.318284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.111782Z digest=sha256:fe99150add6881eda5b3bb07b45a7104be0da720ef8b7e435e486f7aa40eeed7

Observation ddccf2aa-fed1-4b33-ada8-97c587bf4b0e · outbound

This paper cites Integrated architectures for learning, planning, and reacting based on approximating dynamic programming.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Integrated architectures for learning, planning, and reacting based on approximating dynamic programming

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.805994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.116205Z digest=sha256:723dfa1afd8737c3ca6593fb8900c8430d6cccd3a374285d76d430acfb4fa249

Observation 4e01e6f2-e993-479d-9e2c-a3a130f50738 · outbound

This paper cites Sutton and Andrew G.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Sutton and Andrew G

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.792552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.120492Z digest=sha256:0c4ac3851189b0e0e71d4275aeda7082f5bfe3f062c03ebd13632f1861351d66

Observation 6f15bc57-ea9e-49ba-816d-260ba7d40a00 · outbound

This paper cites DeepMind Control Suite.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization DeepMind Control Suite

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.124594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.124594Z digest=sha256:3db8a222d6d802a614cd5c7b30243666aefb1233d2fd0ecbe2aff90771e41002

Observation 229608f7-4f96-43a3-b34c-32edd1ff9e11 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Mujoco: A physics engine for model-based control

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.778523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.129108Z digest=sha256:c934d568a343291e784fd9ed3373d609bf26fe58781762ddceb86a271860a083

Observation e0aa07f3-b903-4b0b-a8ee-26b53270196f · outbound

This paper cites When to use parametric models in reinforcement learning? In Advances in Neural Information Processing Systems, 2019.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization When to use parametric models in reinforcement learning? In Advances in Neural Information Processing Systems, 2019

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.765022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.133207Z digest=sha256:8c92784f58fbf843fc94e9f8ac8ff530fc93f61a7359eab44b13494ca2c059c6

Observation caed7d6c-b0b2-48a8-8a35-6c0e6eaba9a9 · outbound

This paper cites L2 Regularization versus Batch and Weight Normalization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization L2 Regularization versus Batch and Weight Normalization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.137727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.137727Z digest=sha256:3d1ee1ce7f2049ab7fef56bcd648b52062d331a2e060ef4fc64abc952ae038ad

Observation e034f50a-2aff-4ba1-806d-58df997a9bee · outbound

This paper cites MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.142069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.142069Z digest=sha256:1a26a5143591040fd7c288433d1a09d9a65ede7dc1f89940bf66ad61dd42be6b

Observation 6c2e5cfb-aa5d-4209-abb3-c3fe1571c58a · outbound

This paper cites Root mean square layer normalization.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Root mean square layer normalization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.751348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.146711Z digest=sha256:df5760e6e0e2dc6b0cc7c06c1b31b05c4252acf045a428a2bb18f11b0ebdc7a9

Observation 9174a6a3-46b7-43ed-9071-aad4835e0467 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.738107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.151181Z digest=sha256:3af4afde376a70b982679a479f3c6659d281f0cf103f457287ef650556ed7ac3

Observation 714c4790-4185-46eb-9328-c49b541c7c85 · outbound

This paper cites Limitations.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Limitations

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.724724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.155477Z digest=sha256:a771bc2ccdae93fcb87231dce68c273896e6996e30e05bb37eaff81dbe6f1bf9

Observation b3af780f-3a10-47f8-80fd-49a647ddfce4 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.711119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.160017Z digest=sha256:418d532b6023c67793a11026b2c01d3e4924dd5e4f7c5634ba9a0055b5516afb

Observation c4887074-6461-48b7-b3bb-c3d71997bd38 · outbound

This paper cites To aid reproducibility, we plan to release the code together with the camera-ready version of the paper.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization To aid reproducibility, we plan to release the code together with the camera-ready version of the paper

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.697899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.164340Z digest=sha256:27397aba7f0aa2f862f448097f3277df0697677dd1b2438eced6e32e90f23599

Observation b674ce1f-a0e4-425e-9a69-476838c88c3c · outbound

This paper cites We plan to release the code together with the publication of the paper.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization We plan to release the code together with the publication of the paper

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.683917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.169410Z digest=sha256:c0695fdf18e838c134ac0f51ae0fa709baefff71e099e2d148c960dd0e2ff794

Observation fb87d9b0-aa13-4434-8428-bb10e92b1ccf · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not include experiments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.670334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.174165Z digest=sha256:6486a34840e05457f8b8bdc8fa9c88fa641cb0628bcf57ef36e7a4efe1ccedec

Observation 79634b99-c3fb-4dd3-bd36-d4f912570c3e · outbound

This paper cites Results are aggregated over multiple environments and 10 seeds each.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Results are aggregated over multiple environments and 10 seeds each

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.656920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.178258Z digest=sha256:0876b86810daabc5b5cd03b0e4c9f53e4528cb638a357ac3e4e96193d534f677

Observation 81b417fa-0671-4dc4-aaa7-ce5973b1c6a6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not include experiments

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.642686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.182482Z digest=sha256:9ff2a3cd692d34ed7e37436b9c10d26ed7133d6b82283a16d85b2ff7925258bd

Observation b74d57b4-a312-4725-91d9-032675b7d9cb · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.627524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.187026Z digest=sha256:d5a96af5a78d4e91e736ab765c9992868172f6cb52c79df9009c0609eb4deeb5

Observation 73ed4cfe-9dba-41ea-8e2c-66046cbf7835 · outbound

This paper cites As actor-critic methods already enjoy a long history, there is no additional societal impact with this research contribution.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization As actor-critic methods already enjoy a long history, there is no additional societal impact with this research contribution

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.613570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.191284Z digest=sha256:70fe2b8e283cdc42c0e8fa9abce5b0c648a4c7d10baf98335102f9258df645d9

Observation fbd4819c-cde6-425b-a14a-2103608adc12 · outbound

This paper cites an unresolved cited work.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:32:20.599457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.195476Z digest=sha256:708d30dc8018272a7a9b26963770f9f26035838938aab8b1b6ef4b390e35d4c6

Observation 1cc4494b-c4ec-49df-94cb-8727302ce406 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not use existing assets

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.585362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.199803Z digest=sha256:22e8e232ea18e836a68caffb73ffabb14ee99c8695301e8e0b0f4f01fe71bcec

Observation 1659f625-38a3-477a-8cb6-ec380efaf8b4 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Guidelines: • The answer NA means that the paper does not release new assets

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.571135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.204221Z digest=sha256:f8d2037f16fcf42fc12adcf63bcde007839b8ebcd985009c1aa09bce72e5316c

Observation 58009948-31b4-4bdf-81cf-8cdde5f88573 · outbound

This paper cites an unresolved cited work.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.208772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.208772Z digest=sha256:51481f2896b8c5e90dd92c5306a97a429b8f3dfecd999143197f078c55756554

Observation c05b1cf6-fef2-4e2e-807c-e4aa83e6e252 · outbound

This paper cites • Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization • Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:32:20.548252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.212997Z digest=sha256:5ec8dd20a7d744f54b813f2577f679f48a489193ba11b9344c84826e0adc7efb

Observation 7a401270-6cba-409f-870d-c645450b258a · outbound

This paper cites an unresolved cited work.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:32:20.533194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:32:20.217328Z digest=sha256:e86695e35453bd8e56e047d73160551201f8a53591a8d37841127ab8a1b68a66

Observation f2f1f8d1-f3ba-4ee3-a3e9-b0a4e6212f51 · outbound

This paper cites Bigger, Better, Faster: Human-level Atari with human-level efficiency.

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization Bigger, Better, Faster: Human-level Atari with human-level efficiency

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:20.107232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:32:20.107232Z digest=sha256:1f70b59a6351d711008ef99e51a796ae1361fe94f793e7298a7021f64a743c8b

Pith citing papers

Observation 47bb1096-30d7-43a5-a96a-02dbf597181b · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.014406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.014406Z digest=sha256:ff04d4d7257fb0d4d0112c85ba45a443a552a6c4615a91493954e34e945058ca

Observation 8e5df350-6c5b-40a5-8584-b9664a6ecbbb · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:27.863900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:ed26c9d9bd0a6b2a3c18233bd7991d16eee099bc76243451811dddacc9c65471

Observation a3db05dd-9295-402a-bc36-ca5c22c7c0fe · inbound

Extending Differential Temporal Difference Methods for Episodic Problems cites this paper.

Extending Differential Temporal Difference Methods for Episodic Problems Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:38.842202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:19:57.472765Z digest=sha256:70b99d95a907d37ffbea76ad47b7cdd32ec9fbfc44de78a457bac1340c7f2177