Pith. sign in

Paper Citation Record · LEDGER

Value-Based Deep RL Scales Predictably

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2502.04327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04327 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:50:34.155640Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:52.599976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:09:12.714863Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a94a2006-4cda-4f14-981a-d5d2dda951c6 · outbound

This paper cites write newline.

Value-Based Deep RL Scales Predictably write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:33.979227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:33.979227Z digest=sha256:997a0b82e427e47d385ec511fb26d05476e4ef4f4d56fa6d092755a1e63e415d

Observation ab0650aa-07fb-499b-9796-cde2ef210b24 · outbound

This paper cites GPT -4 technical report.

Value-Based Deep RL Scales Predictably GPT -4 technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.673980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.983471Z digest=sha256:3feb91eaf9816d69404063ea113c15bcfe1e8d4013391b1bfd96360ba3d9dc77

Observation 46769471-4517-4647-ad00-03ab6025acd0 · outbound

This paper cites The isotonic regression problem and its dual.

Value-Based Deep RL Scales Predictably The isotonic regression problem and its dual

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.664945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.986597Z digest=sha256:5c47fc83a3e0450643f2ede9e3414c9a3993d88dc5525311c041781e0ab83590

Observation d089a7e0-efd8-4d33-834b-0c01e5af1a3a · outbound

This paper cites Pattern Recognition and Machine Learning.

Value-Based Deep RL Scales Predictably Pattern Recognition and Machine Learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.655527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.990214Z digest=sha256:50121e563ac98d72e75ca43362f2534bb2343f82f98423461ce3132a49efe571

Observation 2a04ab3d-9f6b-422c-a0dc-1c5360bd21ee · outbound

This paper cites OpenAI Gym , 2016.

Value-Based Deep RL Scales Predictably OpenAI Gym , 2016

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.646732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.993346Z digest=sha256:0533b42b385b11e1828d10ffbe1d928681b3e644029a5ab0273e24a035ea2576

Observation 7eb90764-31b0-40c4-b454-096967c9bbc5 · outbound

This paper cites Video generation models as world simulators.

Value-Based Deep RL Scales Predictably Video generation models as world simulators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:33.996307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:33.996307Z digest=sha256:d043f169b0a96548b90c9cd39a8aad380a3530e1bdc2965bbc6742eaf843d88f

Observation 852f239f-9b6c-4765-8c63-6b52b164378d · outbound

This paper cites Randomized ensembled double Q -learning: Learning fast without a model.

Value-Based Deep RL Scales Predictably Randomized ensembled double Q -learning: Learning fast without a model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.632177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.999739Z digest=sha256:7e0f4b56bdfb13cd1663218209988879cfb8754b6e2355f1c2f389e47c7673b6

Observation 60159157-8014-4dbf-88d1-249abf341791 · outbound

This paper cites The value-improvement path: Towards better representations for reinforcement learning.

Value-Based Deep RL Scales Predictably The value-improvement path: Towards better representations for reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.623340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.002855Z digest=sha256:f83c5c31ed0a5f61c6991f94864c753db44a70f9f045bfa873c89b4d6c9817ec

Observation 3f4817d9-e105-4c53-9a66-0ca15e87e663 · outbound

This paper cites Sample-efficient reinforcement learning by breaking the replay ratio barrier.

Value-Based Deep RL Scales Predictably Sample-efficient reinforcement learning by breaking the replay ratio barrier

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.614096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.005965Z digest=sha256:c291b6bbff76367b705141f8f11d0b8a8d69bfea9b5f8b5bb1de883729258c90

Observation b2c7db4c-7743-4e74-9ef8-232c268208c6 · outbound

This paper cites The Llama 3 herd of models.

Value-Based Deep RL Scales Predictably The Llama 3 herd of models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.605256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.009020Z digest=sha256:de85ee9447f0e8e6bcd2b7080331165ac8ce84eaaf8bcac0b39c3792df9e8c03

Observation 8cc2861c-5e2b-4a33-b119-9338a2ee829a · outbound

This paper cites IMPALA : Scalable distributed deep- RL with importance weighted actor-learner architectures.

Value-Based Deep RL Scales Predictably IMPALA : Scalable distributed deep- RL with importance weighted actor-learner architectures

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.596658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.011583Z digest=sha256:3742a94107e1f70e41d7f1a7a326174bdd9c7160b9afc92237ee6a9e9989f900

Observation d5250699-4756-407d-8bbb-dc4baada40de · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Value-Based Deep RL Scales Predictably Language models scale reliably with over-training and on downstream tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.587849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.014741Z digest=sha256:c722070ff5317f80d8ef3a3bf103e9f2221c230d8f8bc7f0163ec33957d766f5

Observation 8ccf7e35-19a6-4457-8bca-a0d23914458b · outbound

This paper cites Simplifying deep temporal difference learning.

Value-Based Deep RL Scales Predictably Simplifying deep temporal difference learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.578847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.017377Z digest=sha256:d9fa3b45e4e33c0c7f341a09f14130e9dab102f6932bda51b472dc7ef3b3dc36

Observation 65dda2fc-cb8a-4df9-831e-6e5e1a2a50b8 · outbound

This paper cites Scaling laws for reward model overoptimization.

Value-Based Deep RL Scales Predictably Scaling laws for reward model overoptimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.020093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.020093Z digest=sha256:29670abda6de977a8ec23e181e3874576c512bf0b1650d2045a6da782d94b70c

Observation 6868a284-8f1c-40c2-b8f4-abef9ae465be · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Value-Based Deep RL Scales Predictably Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.564758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.022820Z digest=sha256:5f92f7ac361799d4660b0fb009b927aa8dd651b3db772f4a14459f771553abe7

Observation a9aabe5d-25c1-4a34-a9d0-7b6e22123557 · outbound

This paper cites Mastering diverse domains through world models.

Value-Based Deep RL Scales Predictably Mastering diverse domains through world models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.556484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.025588Z digest=sha256:fff1f0748d943788e444e293e2e1dd0754e2ff85bb7bb0b7460a7e7c396636aa

Observation f93fd152-ee35-4bb2-8b61-7b5df227e8e4 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

Value-Based Deep RL Scales Predictably Scaling laws for single-agent reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.547912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.028174Z digest=sha256:3e2dc2d16c8127b372910bbf1483fc49bb84468fd79c0b7a0887d190d0a1e147

Observation 38c12941-3766-4884-abc6-5dd8d4ade8a0 · outbound

This paper cites Training compute-optimal large language models.

Value-Based Deep RL Scales Predictably Training compute-optimal large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.539244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.031060Z digest=sha256:90ff4ba98fe87e1a5dd8f8873724df47d8272fe04b67c89e8c9e35df1e53871a

Observation d9a6a78b-3850-451b-a8f8-8d996de70a2f · outbound

This paper cites When to trust your model: Model-based policy optimization.

Value-Based Deep RL Scales Predictably When to trust your model: Model-based policy optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.530725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.033876Z digest=sha256:dbe9c9f155498a8ef4c00df4a0c077b7ea8ede7b417dd0704ff780eaab184523

Observation 445f83e8-9890-4e60-9939-858eafe69d29 · outbound

This paper cites an unresolved cited work.

Value-Based Deep RL Scales Predictably Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:50:34.522130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.037606Z digest=sha256:05e26414106402a9edfc5069f1d18492b8128818ba3be3585710275d44a705ee

Observation 834dc572-89e9-4b65-8a16-51917ffbe58c · outbound

This paper cites Scaling laws for neural language models.

Value-Based Deep RL Scales Predictably Scaling laws for neural language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.513941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.040612Z digest=sha256:ed79f16796b39d552777dd4b9297492841029dc3e456ce568254057e7f799ce6

Observation 4318d582-b797-4984-95ff-7d4fa607f897 · outbound

This paper cites One weird trick for parallelizing convolutional neural networks.

Value-Based Deep RL Scales Predictably One weird trick for parallelizing convolutional neural networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.505067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.043514Z digest=sha256:8fb443a2913efd4d2de09b62fb63120cdf0e16e496d452cb2cc3194f02fc7b11

Observation 45af0646-657e-42e6-81ce-8e5607ea19d4 · outbound

This paper cites Implicit under-parameterization inhibits data-efficient deep reinforcement learning.

Value-Based Deep RL Scales Predictably Implicit under-parameterization inhibits data-efficient deep reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.496003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.046258Z digest=sha256:6faeab48972d484a587c658c8381e32da0d9ebbe11ed6f7b48ed513fda573b3a

Observation 6d8cd011-d773-4ed7-a2bf-4ba6b7b85091 · outbound

This paper cites DR3 : Value-based deep reinforcement learning requires explicit regularization.

Value-Based Deep RL Scales Predictably DR3 : Value-based deep reinforcement learning requires explicit regularization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.486942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.049833Z digest=sha256:37b59b2b828487f2603ebb908f9edef3f1fc75382db08c779a485775165f4203

Observation 2d7d541c-d0e3-4144-b23a-b5ac511ae60d · outbound

This paper cites Offline Q -learning on diverse multi-task data both scales and generalizes.

Value-Based Deep RL Scales Predictably Offline Q -learning on diverse multi-task data both scales and generalizes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.477670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.052642Z digest=sha256:a893b76d27fd599fda0234de0aca0ecfcdad8d248c5b6f2c34b75aeecd91cdc0

Observation a77f9a6a-8821-42b1-8285-193d73da3323 · outbound

This paper cites Plastic: Improving input and label plasticity for sample efficient reinforcement learning.

Value-Based Deep RL Scales Predictably Plastic: Improving input and label plasticity for sample efficient reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.468705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.055443Z digest=sha256:8d7cb7f2698f4f7f67e0144023ba1716716821a27dd94af66f15053e2245da63

Observation 42100522-0b73-4742-a361-b3e43167bb73 · outbound

This paper cites SimBa : Simplicity bias for scaling up parameters in deep reinforcement learning.

Value-Based Deep RL Scales Predictably SimBa : Simplicity bias for scaling up parameters in deep reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.459855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.058595Z digest=sha256:58e126eb1784a227b53ece9b0179d83c6d1d8cb653f515c3d79f2d772ae44e51

Observation 3aaa09b1-f44e-4579-bef2-a6aa6218ac8a · outbound

This paper cites Offline reinforcement learning: Tutorial, review, and perspectives on open problems.

Value-Based Deep RL Scales Predictably Offline reinforcement learning: Tutorial, review, and perspectives on open problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.061981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.061981Z digest=sha256:928c98f6566be06e2731b76a55aa2c8bf13c57838f42f420b9af961f188f1310

Observation 6f7126d6-f226-4751-9d6f-f5391b2d76eb · outbound

This paper cites Efficient deep reinforcement learning requires regulating overfitting.

Value-Based Deep RL Scales Predictably Efficient deep reinforcement learning requires regulating overfitting

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.445968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.064627Z digest=sha256:795b9807e25c02060b63fcac5338c0dd4771c22abca0db98af3859fbb8eb92a7

Observation 8123446f-328a-45db-ac53-6d0b50893b5b · outbound

This paper cites Parallel Q -learning: Scaling off-policy reinforcement learning under massively parallel simulation.

Value-Based Deep RL Scales Predictably Parallel Q -learning: Scaling off-policy reinforcement learning under massively parallel simulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.436750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.067376Z digest=sha256:0b0bfd9363c588e971bc72ba00c175191887e3d1c1e4d1e3b652f4411526eb33

Observation da99da43-bab3-4570-aadd-b9f5ac7bc6b6 · outbound

This paper cites Continuous control with deep reinforcement learning.

Value-Based Deep RL Scales Predictably Continuous control with deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.427523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.070061Z digest=sha256:326857b64d6fdc51b63072539181ff1e4d0935387fdd86d80b9ae8db29c59350

Observation 84f79e94-a75d-49cf-b86f-7c976fd631aa · outbound

This paper cites Scaling laws for fine-grained mixture of experts.

Value-Based Deep RL Scales Predictably Scaling laws for fine-grained mixture of experts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.418696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.072762Z digest=sha256:0a74a1b88666450db324967107caf1887f7508cc0a27bdaf92699c64f530fea4

Observation 95c8a962-f99a-4380-b627-72f3c86d5438 · outbound

This paper cites Understanding plasticity in neural networks.

Value-Based Deep RL Scales Predictably Understanding plasticity in neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.409854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.075389Z digest=sha256:0dbd2def8ffeef61450904a9bb63e91b40232ae8ca1d80ca40c410d260782fe9

Observation 420e74c0-e1dd-4c9d-bd63-330696925885 · outbound

This paper cites Isaac Gym : High performance GPU -based physics simulation for robot learning.

Value-Based Deep RL Scales Predictably Isaac Gym : High performance GPU -based physics simulation for robot learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.401201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.078156Z digest=sha256:ca89e81da3ea51243353901510b6c5e993d5a58b575735fded8b7a4398d20313

Observation 0d76c8ba-933a-4a8e-8993-da19d8bfa3d0 · outbound

This paper cites An empirical model of large-batch training.

Value-Based Deep RL Scales Predictably An empirical model of large-batch training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.392251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.080820Z digest=sha256:ad3a57b03e2e13835c854bb3e9fa6610cc68454d2f607aead657ff51b959f79d

Observation 48466706-5f65-4e81-9dc2-308b1bd52c5e · outbound

This paper cites Human-level control through deep reinforcement learning.

Value-Based Deep RL Scales Predictably Human-level control through deep reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.083679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.083679Z digest=sha256:487502a821f75020939e15427596da7361aa74f2c7106e217b0a04755db614d5

Observation f3dbbb1e-7b6f-4a92-bbf1-8f16247f1a65 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Value-Based Deep RL Scales Predictably Asynchronous methods for deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.378538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.086479Z digest=sha256:8b92ba2d505aea0d649c9b1da71661bda1a61ec258c22c761bad4d12ebb86faf

Observation ee438f56-4ef8-409c-904c-0dbf3daf1055 · outbound

This paper cites Scaling data-constrained language models.

Value-Based Deep RL Scales Predictably Scaling data-constrained language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.369315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.089206Z digest=sha256:857767398f3e0f1ec467fcf024b02c6e36ee3a6098ffd3bdfc42da36cb03b0e0

Observation 0c005f43-a340-46aa-a5df-b1260383fd20 · outbound

This paper cites Overestimation, overfitting, and plasticity in actor-critic: The bitter lesson of reinforcement learning.

Value-Based Deep RL Scales Predictably Overestimation, overfitting, and plasticity in actor-critic: The bitter lesson of reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.360790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.092049Z digest=sha256:1a92bf771c41af9eb5eca0f7f51cad42902a086fc05d72fb0cc0710905bec9a9

Observation 3da14635-a3a1-4516-bf12-8fb94a1243ee · outbound

This paper cites Bigger, regularized, optimistic: Scaling for compute and sample-efficient continuous control.

Value-Based Deep RL Scales Predictably Bigger, regularized, optimistic: Scaling for compute and sample-efficient continuous control

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.351722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.094759Z digest=sha256:8a8cb52e9301615952c864a98e9ff771cce237a976ce6df167a6c938fde01af0

Observation 15f7c4a0-458e-488e-b1be-109c74191160 · outbound

This paper cites The primacy bias in deep reinforcement learning.

Value-Based Deep RL Scales Predictably The primacy bias in deep reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.342821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.097737Z digest=sha256:6fd4a4bd2136b846e1effff15054b1b64e511b3313542912c61303ed2872b883

Observation 1b2963ce-248d-4e98-8dd4-4600f46c7c7b · outbound

This paper cites Is value learning really the main bottleneck in offline RL ? Advances in Neural Information Processing Systems, 2024.

Value-Based Deep RL Scales Predictably Is value learning really the main bottleneck in offline RL ? Advances in Neural Information Processing Systems, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.333902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.100435Z digest=sha256:ed3d63c3f5537b6b7d358f76ab5b7907c9b5109eb80cedfab4593a2b4845fa64

Observation a1946095-49fa-4c8a-bbef-b856df7aba23 · outbound

This paper cites Hierarchical text-conditional image generation with CLIP latents.

Value-Based Deep RL Scales Predictably Hierarchical text-conditional image generation with CLIP latents

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.325067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.103094Z digest=sha256:ad428e71b2d3cf0ce120bd105a5f5a4e259164868b51a8f9efbe3ea4f022015d

Observation 2e36fc28-0661-4ad9-bcd8-81ada1a772bb · outbound

This paper cites Mastering Atari , Go , chess and Shogi by planning with a learned model.

Value-Based Deep RL Scales Predictably Mastering Atari , Go , chess and Shogi by planning with a learned model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.316118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.106006Z digest=sha256:2782b309e26a607173bf43fe2ad2ea85247c98497e48b095684c55cbfb846c62

Observation 9f4e15e0-c640-431a-9fb8-700b030b5344 · outbound

This paper cites Proximal policy optimization algorithms.

Value-Based Deep RL Scales Predictably Proximal policy optimization algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.108691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.108691Z digest=sha256:ebcb59b2c1fb5ac8d20b9b33336b2ed702c63fe57dbbd8d82bb3dc11bcd5896c

Observation 449bf428-f076-4b1c-94f1-bb0404809cd8 · outbound

This paper cites Bigger, better, faster: Human-level Atari with human-level efficiency.

Value-Based Deep RL Scales Predictably Bigger, better, faster: Human-level Atari with human-level efficiency

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.302091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.111324Z digest=sha256:9f0777321367ca41cf95fd573b6464de27f94113fc9ee5de248b46bb43646f6f

Observation 5a430338-03f3-4ad6-b4c5-a34cd3733eec · outbound

This paper cites Mastering the game of Go with deep neural networks and tree search.

Value-Based Deep RL Scales Predictably Mastering the game of Go with deep neural networks and tree search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.113924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.113924Z digest=sha256:62b2d44b9ca9adb9ca9dde30aad7f72f6ca4484d82e5a94636ebb95d602ec4c5

Observation 3eaa4bac-7c82-40f6-b3df-9b940bd3a4d9 · outbound

This paper cites SAPG : Split and aggregate policy gradients.

Value-Based Deep RL Scales Predictably SAPG : Split and aggregate policy gradients

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.288447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.116676Z digest=sha256:a4f1e2876a5b9a975248f2fe682110902fe97747e64bd0ba1969a9803c9a2d6c

Observation 13a14c9a-0a6f-42c0-a627-5d31545b7602 · outbound

This paper cites The dormant neuron phenomenon in deep reinforcement learning.

Value-Based Deep RL Scales Predictably The dormant neuron phenomenon in deep reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.279987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.119692Z digest=sha256:008d9714b143b3a9dcbff6d962c2e95010709ccd77a84ed781126dcbb262d75e

Observation d880c06a-a724-438d-887c-f8e3e5ce3455 · outbound

This paper cites Offline actor-critic reinforcement learning scales to large models.

Value-Based Deep RL Scales Predictably Offline actor-critic reinforcement learning scales to large models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.271721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.122488Z digest=sha256:1dce18dc40540361f2d75b27a48ea7e1cfe82a15dc1c8a082ad57dff1fea93db

Observation c808fa9a-2d26-4506-a96a-d5b8cf981156 · outbound

This paper cites Reinforcement Learning: An Introduction.

Value-Based Deep RL Scales Predictably Reinforcement Learning: An Introduction

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.263108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.125185Z digest=sha256:0024517b20b6d5fb020c3c2c57884d5984d53a2732dd26090b805317a947ebc9

Observation fde0c6e6-6d02-4df5-af67-dbcdad5eef6b · outbound

This paper cites DeepMind control suite.

Value-Based Deep RL Scales Predictably DeepMind control suite

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.254228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.128913Z digest=sha256:824d1caf2eb80f067e4cd62b6fbc39b0a2b341699501f506cdb674e56ec4f088

Observation 4f5daa81-71cf-47e6-9a55-6a1aa11b9cef · outbound

This paper cites Gemini : A family of highly capable multimodal models.

Value-Based Deep RL Scales Predictably Gemini : A family of highly capable multimodal models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.245209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.131743Z digest=sha256:05e3ffe02bf2e81552a5d9b1ed4fcf3f3a511ae195d039d0e7076309953fc1da

Observation a61064de-eb67-4b09-9730-3f2f821f42e5 · outbound

This paper cites dm\_control: Software and tasks for continuous control.

Value-Based Deep RL Scales Predictably dm\_control: Software and tasks for continuous control

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.236045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.134688Z digest=sha256:8e07081dbec0722be22e47118d23d52c270f025e2c35c3af0f42b83404d007b4

Observation 2e34445f-b5a9-4d6f-afdb-e3684d360775 · outbound

This paper cites SciPy 1.0: Fundamental algorithms for scientific computing in Python.

Value-Based Deep RL Scales Predictably SciPy 1.0: Fundamental algorithms for scientific computing in Python

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.227606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.137501Z digest=sha256:09a9dfef8fafedae9f1d19e3bd64de764bcff4de6d0ae3019d7036740080128c

Observation 4eff357b-0b1b-4dae-9a4b-5508821a0fdc · outbound

This paper cites DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization.

Value-Based Deep RL Scales Predictably DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.140276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.140276Z digest=sha256:087f41a2b045552740dba5108e1a6204af7058eb27b9c3039cc6f4d922cc512b

Observation de0631a8-5a78-4ab3-8c53-81235939c286 · outbound

This paper cites Tensor programs V : Tuning large neural networks via zero-shot hyperparameter transfer.

Value-Based Deep RL Scales Predictably Tensor programs V : Tuning large neural networks via zero-shot hyperparameter transfer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.218727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.143693Z digest=sha256:4e537ca876c606e64baa38b486e5f1f2c75bbb93d5ac8361d6b803142755e7ee

Observation 3c1fa771-e5cd-4d0a-979f-d37b620ecc1a · outbound

This paper cites How to leverage unlabeled data in offline reinforcement learning.

Value-Based Deep RL Scales Predictably How to leverage unlabeled data in offline reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.209081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.146473Z digest=sha256:fa38b32dc7e5bf95a6b28519ca69c5b7e2261c4176de45b6145683099bd334a9

Observation b2c92c87-3905-4123-b6a1-b0bbf2073a2d · outbound

This paper cites @esa (Ref.

Value-Based Deep RL Scales Predictably @esa (Ref

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.149185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.149185Z digest=sha256:382f183c9677cfc43a64bc0149188946a317d6ec4a335aa0b221376e0328daac

Observation a06c7c6a-6d48-47bf-a2db-61889dbf1be1 · outbound

This paper cites an unresolved cited work.

Value-Based Deep RL Scales Predictably Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.152692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.152692Z digest=sha256:a0a1087071217cda9dc7bee014b3b6f695c57de3f55d2769b17644ce214d568c

Observation 5295faeb-a8ea-4c43-854e-c51d9f0ff23c · outbound

This paper cites an unresolved cited work.

Value-Based Deep RL Scales Predictably Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.155640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.155640Z digest=sha256:490aa59a1065e3815e1542327008e3dc20c98dc151672491a739b897e0cb95fc

Pith citing papers

Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · inbound

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners cites this paper.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.599976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.599976Z digest=sha256:4fbb26a76069ca823bf06f9a876764b51aeca4d9dd018a59d288da5c70987e40

Observation e7d746db-8c02-4bce-8d81-72aab31f19d0 · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Value-Based Deep RL Scales Predictably

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.797394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.797394Z digest=sha256:453a93fb5a41636124690f818261a5eb16642ceea68a46755e5d5a012063abe4

Observation b6387637-3a7d-4c8d-ab63-6aa13395de49 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.580539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:33:19.038889Z digest=sha256:756dad42a119ecc60c02980dfff73daf37f1339c6441855b33bb287ab070f813

Observation cadfc5c9-2956-4f39-ade1-1c0075c43668 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.810267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:27:38.643667Z digest=sha256:c9d3a5d88cff2203963e29d1d5d99493eb549c0c84dbaa4460ee27c268dab8c5

Observation 19d61ca9-241b-4520-8130-36a7f28b563c · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:12.718570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:04:12.943222Z digest=sha256:315bc68f4eb6173364dd1d96594a9ed0cadb71d71d192220559a717847a90c20