Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:20:23.254360Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 10 inbound Pith citation observations for arXiv:1908.03568.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:20:23.254360Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T11:15:24.591790Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T07:55:33.101030Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9baa8757-7f4f-4e32-b44f-144137e4f63f · outbound
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad1c6a5-5942-4296-9177-84a3850e4311 · outbound
Behaviour Suite for Reinforcement Learning The gains are most extreme in the exploration tasks, where ensemble sizes less than 10 are not able to solve large ‘deep sea’ tasks, but larger ensembles solve them reliably
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f51aef5a-e9bd-4967-b021-8fccafb0e26b · outbound
Behaviour Suite for Reinforcement Learning Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd825e92-858d-4893-bc24-b253875fdb47 · outbound
Behaviour Suite for Reinforcement Learning Deep Exploration via Randomized Value Functions
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6fb266a-a168-4062-94b2-22308bf15fc2 · outbound
Behaviour Suite for Reinforcement Learning doi: 10.1126/science.aar6404
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2917ceb3-dc54-44e5-bf51-c3889ecbe66b · outbound
Behaviour Suite for Reinforcement Learning Benchmarking Model-Based Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a1203d-1320-4a14-bb38-64e0581af1a1 · outbound
Behaviour Suite for Reinforcement Learning Kingma and Jimmy Ba
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5da487-e568-4ca8-bd3b-a0e00ea15dd3 · outbound
Behaviour Suite for Reinforcement Learning A Tutorial on Thompson Sampling
Reference 1958
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a57811-37d3-4691-9b75-de3f609fa628 · outbound
Behaviour Suite for Reinforcement Learning MIPLIB 2017,
Reference 1961
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8d35d5b1-7b79-4344-b655-8ca6ba06e2e6 · outbound
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68cf11e6-f440-4901-b252-7deaf61de2db · outbound
Behaviour Suite for Reinforcement Learning doi: 10.1109/ MCSE.2007.53
Reference 2007
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eeab421e-c2d9-4e84-9da8-1404c2f0436f · outbound
Behaviour Suite for Reinforcement Learning DeepMind Control Suite
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 270ae6cc-5677-4ca5-a524-b90d898b103c · outbound
Behaviour Suite for Reinforcement Learning Large-scale machine learning with stochastic gradient descent
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 48945cdc-2f9c-451c-8c52-428099f88c79 · outbound
Behaviour Suite for Reinforcement Learning Adam: A Method for Stochastic Optimization
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a79a638-af2f-43a4-8b98-05ea01fa4bde · outbound
Behaviour Suite for Reinforcement Learning Reconciling modern machine learning practice and the bias-variance trade-off
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e9faa9-8ba1-4236-82d7-9c0ef3e8cfaf · outbound
Behaviour Suite for Reinforcement Learning Deep Reinforcement Learning that Matters
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba3e373-2518-4a4d-82c7-e0e053a0e4b1 · outbound
Behaviour Suite for Reinforcement Learning Dopamine: A Research Framework for Deep Reinforcement Learning
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980b90bc-3cbc-4cf6-a236-f6d72422b264 · outbound
Behaviour Suite for Reinforcement Learning Unresolved cited work
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5ed7b8d2-1fe1-46ae-b36b-4e6f8f435c81 · inbound
OpenSpiel: A Framework for Reinforcement Learning in Games Behaviour Suite for Reinforcement Learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9859f81-6607-459b-bcdf-32844123fbe2 · inbound
On the Measure of Intelligence Behaviour Suite for Reinforcement Learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 377a83e6-8a25-40b2-9962-1b9c8b7d8b63 · inbound
Mastering Diverse Domains through World Models Behaviour Suite for Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 951569ed-fd37-43a5-82d5-0f64b2f2ceb3 · inbound
Gymnasium: A Standard Interface for Reinforcement Learning Environments Behaviour Suite for Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a01d860-ebdf-40f6-b019-47e29a6d6f97 · inbound
Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Behaviour Suite for Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5c2754d1-30f3-459f-8ed0-24229f650e6d · inbound
Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Behaviour Suite for Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 68b25ed3-1791-403b-a4c5-2f4cfd4edf95 · inbound
On the Effect of Regularization in Policy Mirror Descent Behaviour Suite for Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 185d5759-6035-4f49-b900-61b64e9109cb · inbound
T-GRAB: A Synthetic Diagnostic Benchmark for Learning on Temporal Graphs Behaviour Suite for Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5957897-8f7f-41ed-977c-ebc1b3adf9fc · inbound
Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX Behaviour Suite for Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e822db9-e1df-4e89-a247-15ccab53491b · inbound
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs Behaviour Suite for Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.