Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:50:34.155640Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2502.04327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:50:34.155640Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:52.599976Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T23:09:12.714863Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a94a2006-4cda-4f14-981a-d5d2dda951c6 · outbound
Value-Based Deep RL Scales Predictably write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0650aa-07fb-499b-9796-cde2ef210b24 · outbound
Value-Based Deep RL Scales Predictably GPT -4 technical report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46769471-4517-4647-ad00-03ab6025acd0 · outbound
Value-Based Deep RL Scales Predictably The isotonic regression problem and its dual
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d089a7e0-efd8-4d33-834b-0c01e5af1a3a · outbound
Value-Based Deep RL Scales Predictably Pattern Recognition and Machine Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a04ab3d-9f6b-422c-a0dc-1c5360bd21ee · outbound
Value-Based Deep RL Scales Predictably OpenAI Gym , 2016
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7eb90764-31b0-40c4-b454-096967c9bbc5 · outbound
Value-Based Deep RL Scales Predictably Video generation models as world simulators
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852f239f-9b6c-4765-8c63-6b52b164378d · outbound
Value-Based Deep RL Scales Predictably Randomized ensembled double Q -learning: Learning fast without a model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 60159157-8014-4dbf-88d1-249abf341791 · outbound
Value-Based Deep RL Scales Predictably The value-improvement path: Towards better representations for reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f4817d9-e105-4c53-9a66-0ca15e87e663 · outbound
Value-Based Deep RL Scales Predictably Sample-efficient reinforcement learning by breaking the replay ratio barrier
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b2c7db4c-7743-4e74-9ef8-232c268208c6 · outbound
Value-Based Deep RL Scales Predictably The Llama 3 herd of models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cc2861c-5e2b-4a33-b119-9338a2ee829a · outbound
Value-Based Deep RL Scales Predictably IMPALA : Scalable distributed deep- RL with importance weighted actor-learner architectures
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d5250699-4756-407d-8bbb-dc4baada40de · outbound
Value-Based Deep RL Scales Predictably Language models scale reliably with over-training and on downstream tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ccf7e35-19a6-4457-8bca-a0d23914458b · outbound
Value-Based Deep RL Scales Predictably Simplifying deep temporal difference learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65dda2fc-cb8a-4df9-831e-6e5e1a2a50b8 · outbound
Value-Based Deep RL Scales Predictably Scaling laws for reward model overoptimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6868a284-8f1c-40c2-b8f4-abef9ae465be · outbound
Value-Based Deep RL Scales Predictably Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a9aabe5d-25c1-4a34-a9d0-7b6e22123557 · outbound
Value-Based Deep RL Scales Predictably Mastering diverse domains through world models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f93fd152-ee35-4bb2-8b61-7b5df227e8e4 · outbound
Value-Based Deep RL Scales Predictably Scaling laws for single-agent reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 38c12941-3766-4884-abc6-5dd8d4ade8a0 · outbound
Value-Based Deep RL Scales Predictably Training compute-optimal large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9a6a78b-3850-451b-a8f8-8d996de70a2f · outbound
Value-Based Deep RL Scales Predictably When to trust your model: Model-based policy optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 445f83e8-9890-4e60-9939-858eafe69d29 · outbound
Value-Based Deep RL Scales Predictably Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 834dc572-89e9-4b65-8a16-51917ffbe58c · outbound
Value-Based Deep RL Scales Predictably Scaling laws for neural language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4318d582-b797-4984-95ff-7d4fa607f897 · outbound
Value-Based Deep RL Scales Predictably One weird trick for parallelizing convolutional neural networks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45af0646-657e-42e6-81ce-8e5607ea19d4 · outbound
Value-Based Deep RL Scales Predictably Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d8cd011-d773-4ed7-a2bf-4ba6b7b85091 · outbound
Value-Based Deep RL Scales Predictably DR3 : Value-based deep reinforcement learning requires explicit regularization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d7d541c-d0e3-4144-b23a-b5ac511ae60d · outbound
Value-Based Deep RL Scales Predictably Offline Q -learning on diverse multi-task data both scales and generalizes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a77f9a6a-8821-42b1-8285-193d73da3323 · outbound
Value-Based Deep RL Scales Predictably Plastic: Improving input and label plasticity for sample efficient reinforcement learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42100522-0b73-4742-a361-b3e43167bb73 · outbound
Value-Based Deep RL Scales Predictably SimBa : Simplicity bias for scaling up parameters in deep reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3aaa09b1-f44e-4579-bef2-a6aa6218ac8a · outbound
Value-Based Deep RL Scales Predictably Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7126d6-f226-4751-9d6f-f5391b2d76eb · outbound
Value-Based Deep RL Scales Predictably Efficient deep reinforcement learning requires regulating overfitting
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8123446f-328a-45db-ac53-6d0b50893b5b · outbound
Value-Based Deep RL Scales Predictably Parallel Q -learning: Scaling off-policy reinforcement learning under massively parallel simulation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation da99da43-bab3-4570-aadd-b9f5ac7bc6b6 · outbound
Value-Based Deep RL Scales Predictably Continuous control with deep reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 84f79e94-a75d-49cf-b86f-7c976fd631aa · outbound
Value-Based Deep RL Scales Predictably Scaling laws for fine-grained mixture of experts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95c8a962-f99a-4380-b627-72f3c86d5438 · outbound
Value-Based Deep RL Scales Predictably Understanding plasticity in neural networks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 420e74c0-e1dd-4c9d-bd63-330696925885 · outbound
Value-Based Deep RL Scales Predictably Isaac Gym : High performance GPU -based physics simulation for robot learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d76c8ba-933a-4a8e-8993-da19d8bfa3d0 · outbound
Value-Based Deep RL Scales Predictably An empirical model of large-batch training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48466706-5f65-4e81-9dc2-308b1bd52c5e · outbound
Value-Based Deep RL Scales Predictably Human-level control through deep reinforcement learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3dbbb1e-7b6f-4a92-bbf1-8f16247f1a65 · outbound
Value-Based Deep RL Scales Predictably Asynchronous methods for deep reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ee438f56-4ef8-409c-904c-0dbf3daf1055 · outbound
Value-Based Deep RL Scales Predictably Scaling data-constrained language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c005f43-a340-46aa-a5df-b1260383fd20 · outbound
Value-Based Deep RL Scales Predictably Overestimation, overfitting, and plasticity in actor-critic: The bitter lesson of reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3da14635-a3a1-4516-bf12-8fb94a1243ee · outbound
Value-Based Deep RL Scales Predictably Bigger, regularized, optimistic: Scaling for compute and sample-efficient continuous control
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 15f7c4a0-458e-488e-b1be-109c74191160 · outbound
Value-Based Deep RL Scales Predictably The primacy bias in deep reinforcement learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b2963ce-248d-4e98-8dd4-4600f46c7c7b · outbound
Value-Based Deep RL Scales Predictably Is value learning really the main bottleneck in offline RL ? Advances in Neural Information Processing Systems, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1946095-49fa-4c8a-bbef-b856df7aba23 · outbound
Value-Based Deep RL Scales Predictably Hierarchical text-conditional image generation with CLIP latents
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e36fc28-0661-4ad9-bcd8-81ada1a772bb · outbound
Value-Based Deep RL Scales Predictably Mastering Atari , Go , chess and Shogi by planning with a learned model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f4e15e0-c640-431a-9fb8-700b030b5344 · outbound
Value-Based Deep RL Scales Predictably Proximal policy optimization algorithms
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449bf428-f076-4b1c-94f1-bb0404809cd8 · outbound
Value-Based Deep RL Scales Predictably Bigger, better, faster: Human-level Atari with human-level efficiency
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a430338-03f3-4ad6-b4c5-a34cd3733eec · outbound
Value-Based Deep RL Scales Predictably Mastering the game of Go with deep neural networks and tree search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eaa4bac-7c82-40f6-b3df-9b940bd3a4d9 · outbound
Value-Based Deep RL Scales Predictably SAPG : Split and aggregate policy gradients
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 13a14c9a-0a6f-42c0-a627-5d31545b7602 · outbound
Value-Based Deep RL Scales Predictably The dormant neuron phenomenon in deep reinforcement learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d880c06a-a724-438d-887c-f8e3e5ce3455 · outbound
Value-Based Deep RL Scales Predictably Offline actor-critic reinforcement learning scales to large models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c808fa9a-2d26-4506-a96a-d5b8cf981156 · outbound
Value-Based Deep RL Scales Predictably Reinforcement Learning: An Introduction
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fde0c6e6-6d02-4df5-af67-dbcdad5eef6b · outbound
Value-Based Deep RL Scales Predictably DeepMind control suite
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f5daa81-71cf-47e6-9a55-6a1aa11b9cef · outbound
Value-Based Deep RL Scales Predictably Gemini : A family of highly capable multimodal models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a61064de-eb67-4b09-9730-3f2f821f42e5 · outbound
Value-Based Deep RL Scales Predictably dm\_control: Software and tasks for continuous control
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e34445f-b5a9-4d6f-afdb-e3684d360775 · outbound
Value-Based Deep RL Scales Predictably SciPy 1.0: Fundamental algorithms for scientific computing in Python
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4eff357b-0b1b-4dae-9a4b-5508821a0fdc · outbound
Value-Based Deep RL Scales Predictably DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0631a8-5a78-4ab3-8c53-81235939c286 · outbound
Value-Based Deep RL Scales Predictably Tensor programs V : Tuning large neural networks via zero-shot hyperparameter transfer
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c1fa771-e5cd-4d0a-979f-d37b620ecc1a · outbound
Value-Based Deep RL Scales Predictably How to leverage unlabeled data in offline reinforcement learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b2c92c87-3905-4123-b6a1-b0bbf2073a2d · outbound
Value-Based Deep RL Scales Predictably @esa (Ref
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06c7c6a-6d48-47bf-a2db-61889dbf1be1 · outbound
Value-Based Deep RL Scales Predictably Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5295faeb-a8ea-4c43-854e-c51d9f0ff23c · outbound
Value-Based Deep RL Scales Predictably Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · inbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d746db-8c02-4bce-8d81-72aab31f19d0 · inbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Value-Based Deep RL Scales Predictably
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6387637-3a7d-4c8d-ab63-6aa13395de49 · inbound
When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cadfc5c9-2956-4f39-ade1-1c0075c43668 · inbound
When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 19d61ca9-241b-4520-8130-36a7f28b563c · inbound
When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.