Pith. sign in

Paper Citation Record · LEDGER

CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:1902.05605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1902.05605 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:11:39.516466Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T12:35:49.046229Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 148ed036-5575-478d-9161-dbfcf92f9950 · inbound

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog cites this paper.

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T12:35:49.049370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T12:32:09.940698Z digest=sha256:6865d7744873f8eb6b4556e96957384d025b4a1ab2a3c8661ae297290b2e9f7e

Observation 6d558e1a-6383-43cb-bb04-6ab26cb2f257 · inbound

Latent Action Learning Requires Supervision in the Presence of Distractors cites this paper.

Latent Action Learning Requires Supervision in the Presence of Distractors CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:18:44.870942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:18:44.870942Z digest=sha256:874d9864e9b11a23dcc609ac204653cc11e53c58c143a9b30581d8c64ecf8119

Observation 86a7f4c2-f730-4022-a8ea-987ca76ab9e6 · inbound

Hadamax Encoding: Elevating Performance in Model-Free Atari cites this paper.

Hadamax Encoding: Elevating Performance in Model-Free Atari CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.587512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.587512Z digest=sha256:c970140aa2103d9419cb162bdc9f83406ba56e4a70b234a4b0fcfaabd4029ef2

Observation 86a6e4d0-061c-4591-9a1b-8bbe15e8a26a · inbound

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments cites this paper.

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:59.140424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:59.140424Z digest=sha256:43ae19d337f5dcb72ca1170d8f3ce3c239277ae791dbd332f254b3db483b4a13

Observation 117da26a-037c-44ad-b26e-7ba98a2bfc67 · inbound

Tactile MNIST: Benchmarking Active Tactile Perception cites this paper.

Tactile MNIST: Benchmarking Active Tactile Perception CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:33.483925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:33.483925Z digest=sha256:e25e9b15208984febd6a1ec48cca3866c014550d5ebd09d525425e685af470d2

Observation 26523690-016d-4f4b-9dc7-d282437ea02a · inbound

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning cites this paper.

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T19:12:08.391612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:12:08.391612Z digest=sha256:9d258606e5e0d97b771fb93687629d51f083a33c6343362a9d1de7be5bae0ebf

Observation 2a1b755e-7fb3-4ed6-92a5-4092499cd511 · inbound

FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control cites this paper.

FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:20:00.681688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T12:17:59.392158Z digest=sha256:11baf20066ee9b592f1452a64d974e937c3d771d7849de14c279913d93a9d01d

Observation 7108ded1-49ab-4188-a104-ff0323b720a5 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.868265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:362344637c75e9c9384342fc29a5f0fee352d4590c709bb6245b4545ff5cea03

Observation 44eba5d9-dc3b-4dfd-9133-5f4843611329 · inbound

AdamO: A Collapse-Suppressed Optimizer for Offline RL cites this paper.

AdamO: A Collapse-Suppressed Optimizer for Offline RL CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:31:03.438730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T14:48:05.166453Z digest=sha256:dd8dce6df1267533bcb0d0d0c9c099505332641b7ea45dd6fa3100d17b09c067

Observation a9aed3e1-703b-4f9f-b7f4-13b961551d02 · inbound

Relative Value Learning cites this paper.

Relative Value Learning CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:32:00.068959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:32:00.068959Z digest=sha256:d95ea4e2f12e5e8b37ff9da6332de90c1b75dc0711433605b136434a7f67cabe

Observation 032e212a-2477-4bde-a4d9-c0a0ff21ca18 · inbound

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks cites this paper.

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T10:11:39.516466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T10:11:39.516466Z digest=sha256:e80ed44ced1d77ca0924e0d9b2e842ae5e344bffaffc5a060fc95393ddbc7879