Pith. sign in

Paper Citation Record · LEDGER

Behaviour Suite for Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 10 inbound Pith citation observations for arXiv:1908.03568.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03568 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:20:23.254360Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:15:24.591790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T07:55:33.101030Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved13
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9baa8757-7f4f-4e32-b44f-144137e4f63f · outbound

This paper cites OpenAI Gym.

Behaviour Suite for Reinforcement Learning OpenAI Gym

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.190946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.190946Z digest=sha256:0e9a4ef25b227b2b5df955601f2e3ec3a76992860385ec4e961f9e2845a5abc4

Observation dad1c6a5-5942-4296-9177-84a3850e4311 · outbound

This paper cites The gains are most extreme in the exploration tasks, where ensemble sizes less than 10 are not able to solve large ‘deep sea’ tasks, but larger ensembles solve them reliably.

Behaviour Suite for Reinforcement Learning The gains are most extreme in the exploration tasks, where ensemble sizes less than 10 are not able to solve large ‘deep sea’ tasks, but larger ensembles solve them reliably

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:20:23.396278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T14:20:23.254360Z digest=sha256:2998eb52d2a0c748104facae2c61500d5118482a4fc2a15f5855e528a5718c60

Observation f51aef5a-e9bd-4967-b021-8fccafb0e26b · outbound

This paper cites Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents.

Behaviour Suite for Reinforcement Learning Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.214929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.214929Z digest=sha256:0defde1f4798d23a04038329a564f6a3dd6b460f9c69dd2749605a995f0c6364

Observation fd825e92-858d-4893-bc24-b253875fdb47 · outbound

This paper cites Deep Exploration via Randomized Value Functions.

Behaviour Suite for Reinforcement Learning Deep Exploration via Randomized Value Functions

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:20:23.325375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T14:20:23.222747Z digest=sha256:026fc37ce827684eba2bd9dc30625ac06a9829e2c4deaefd4b59b3aa9f83ab33

Observation e6fb266a-a168-4062-94b2-22308bf15fc2 · outbound

This paper cites doi: 10.1126/science.aar6404.

Behaviour Suite for Reinforcement Learning doi: 10.1126/science.aar6404

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.239196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.239196Z digest=sha256:7a65e46aa3864aa9a33986432970c9f5c072cd039e4d52d6e0ad68bab601a39c

Observation 2917ceb3-dc54-44e5-bf51-c3889ecbe66b · outbound

This paper cites Benchmarking Model-Based Reinforcement Learning.

Behaviour Suite for Reinforcement Learning Benchmarking Model-Based Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.249856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.249856Z digest=sha256:cf7a8545f0ec88bdba71b0f2180a85ab559df45458a48d7a20940a40b5783b0d

Observation e1a1203d-1320-4a14-bb38-64e0581af1a1 · outbound

This paper cites Kingma and Jimmy Ba.

Behaviour Suite for Reinforcement Learning Kingma and Jimmy Ba

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.207094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.207094Z digest=sha256:5aaf66f4267a5de09760c9ebd071d840f9af760bb73b09f0177642a65e7fb57a

Observation ab5da487-e568-4ca8-bd3b-a0e00ea15dd3 · outbound

This paper cites A Tutorial on Thompson Sampling.

Behaviour Suite for Reinforcement Learning A Tutorial on Thompson Sampling

Reference 1958

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.235177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.235177Z digest=sha256:e2fec19c024accab4907fa1aa14220f28970f0f711c2cfb43b788bd6db6a25a8

Observation e9a57811-37d3-4691-9b75-de3f609fa628 · outbound

This paper cites MIPLIB 2017,.

Behaviour Suite for Reinforcement Learning MIPLIB 2017,

Reference 1961

Resolution
parse uncertain
raw_fallback, observed 2026-08-14T14:20:23.430030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T14:20:23.219154Z digest=sha256:cca81ffa60280fb85d565712242ff69db23f3f3817690046d7026df20597cc74

Observation 8d35d5b1-7b79-4344-b655-8ca6ba06e2e6 · outbound

This paper cites DeepMind Lab.

Behaviour Suite for Reinforcement Learning DeepMind Lab

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.172599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.172599Z digest=sha256:b0cbaed16a2e1ef6cc348a6d46b1bc28ad03b4a83420a3396bb9b8e6f7c78465

Observation 68cf11e6-f440-4901-b252-7deaf61de2db · outbound

This paper cites doi: 10.1109/ MCSE.2007.53.

Behaviour Suite for Reinforcement Learning doi: 10.1109/ MCSE.2007.53

Reference 2007

Resolution
malformed identifier
raw_fallback, observed 2026-08-14T14:20:23.406857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T14:20:23.231779Z digest=sha256:fecba1575e7204d806406ea4f729e0d44bf0130a858d75630868e3830affbd0f

Observation eeab421e-c2d9-4e84-9da8-1404c2f0436f · outbound

This paper cites DeepMind Control Suite.

Behaviour Suite for Reinforcement Learning DeepMind Control Suite

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.242776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.242776Z digest=sha256:022b38ff250217283bc9ff80a2e75c2bec674167558ec0b1f204b74a9a1e4ab0

Observation 270ae6cc-5677-4ca5-a524-b90d898b103c · outbound

This paper cites Large-scale machine learning with stochastic gradient descent.

Behaviour Suite for Reinforcement Learning Large-scale machine learning with stochastic gradient descent

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:20:23.446386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T14:20:23.182585Z digest=sha256:c610c1421dffe62845bcf4c14af4e09387d2abed00d70fb82cb5f4ea1b33d6da

Observation 48945cdc-2f9c-451c-8c52-428099f88c79 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Behaviour Suite for Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.210769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.210769Z digest=sha256:4a19e7f3270d9f5d53858e819ddb3470f42cd8ce4e56acc3bad2febc3ccef093

Observation 4a79a638-af2f-43a4-8b98-05ea01fa4bde · outbound

This paper cites Reconciling modern machine learning practice and the bias-variance trade-off.

Behaviour Suite for Reinforcement Learning Reconciling modern machine learning practice and the bias-variance trade-off

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.177818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.177818Z digest=sha256:9be14b0a9a6a9542b46e9467e2f5ba88c48dfb671fd05cc389d30f53be2b76a4

Observation 12e9faa9-8ba1-4236-82d7-9c0ef3e8cfaf · outbound

This paper cites Deep Reinforcement Learning that Matters.

Behaviour Suite for Reinforcement Learning Deep Reinforcement Learning that Matters

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.202658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.202658Z digest=sha256:2259d0dda2c517f8e34075ccfdcde55f5b3aa8d06015832dfbfc7cc43be6e877

Observation 3ba3e373-2518-4a4d-82c7-e0e053a0e4b1 · outbound

This paper cites Dopamine: A Research Framework for Deep Reinforcement Learning.

Behaviour Suite for Reinforcement Learning Dopamine: A Research Framework for Deep Reinforcement Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T14:20:23.194486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:20:23.194486Z digest=sha256:38e87a04b0845181c50c12ae136decd5c1422d3d25104033e1ed0864b5c2885f

Observation 980b90bc-3cbc-4cf6-a236-f6d72422b264 · outbound

This paper cites an unresolved cited work.

Behaviour Suite for Reinforcement Learning Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:20:23.419055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T14:20:23.227474Z digest=sha256:927461120ff5c454882c6e961456cd5404d3ae82e01d3e830aea2eaa81c7fe06

Pith citing papers

Observation 5ed7b8d2-1fe1-46ae-b36b-4e6f8f435c81 · inbound

OpenSpiel: A Framework for Reinforcement Learning in Games cites this paper.

OpenSpiel: A Framework for Reinforcement Learning in Games Behaviour Suite for Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-14T11:15:24.591790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:15:24.591790Z digest=sha256:490f1515c886e0e91656c67174e0bfaaed8b974f18ffdac2923ec4f3d24c4d7e

Observation d9859f81-6607-459b-bcdf-32844123fbe2 · inbound

On the Measure of Intelligence cites this paper.

On the Measure of Intelligence Behaviour Suite for Reinforcement Learning

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:05:34.208295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T13:05:34.106110Z digest=sha256:a057f7c6fad698b8b4ddc7da4e86376961f565f82fdce5c6e513428a27d7077e

Observation 377a83e6-8a25-40b2-9962-1b9c8b7d8b63 · inbound

Mastering Diverse Domains through World Models cites this paper.

Mastering Diverse Domains through World Models Behaviour Suite for Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:08:22.279884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T09:08:21.677362Z digest=sha256:cbdf96b321a757032bcf651d88bcb3e0a7bbe581f72d68fcb28b992282c87e2e

Observation 951569ed-fd37-43a5-82d5-0f64b2f2ceb3 · inbound

Gymnasium: A Standard Interface for Reinforcement Learning Environments cites this paper.

Gymnasium: A Standard Interface for Reinforcement Learning Environments Behaviour Suite for Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:29:49.477736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T17:29:49.186565Z digest=sha256:585076009df48140b979eeb306bb7bfb0acd61d1e178340103f63e424c76eabc

Observation 2a01d860-ebdf-40f6-b019-47e29a6d6f97 · inbound

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior cites this paper.

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Behaviour Suite for Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.399717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T02:38:26.196000Z digest=sha256:f3fe0f18d2789bf8287331b3f3d7da15d54b50e5667e245d6bb4a0fb79438631

Observation 5c2754d1-30f3-459f-8ed0-24229f650e6d · inbound

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior cites this paper.

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Behaviour Suite for Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:55:33.104117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T07:53:48.604436Z digest=sha256:312ffa0b0c585defbf9085a17c6b8a9cedfe65743db68c93216b105ba5b6e274

Observation 68b25ed3-1791-403b-a4c5-2f4cfd4edf95 · inbound

On the Effect of Regularization in Policy Mirror Descent cites this paper.

On the Effect of Regularization in Policy Mirror Descent Behaviour Suite for Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.955783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:18:46.955783Z digest=sha256:ed0f41b75b85eca1290a91b41233ad899a385e4d460f55b794820231f2143062

Observation 185d5759-6035-4f49-b900-61b64e9109cb · inbound

T-GRAB: A Synthetic Diagnostic Benchmark for Learning on Temporal Graphs cites this paper.

T-GRAB: A Synthetic Diagnostic Benchmark for Learning on Temporal Graphs Behaviour Suite for Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:34.092700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:34.092700Z digest=sha256:11d09fd090fbaa260f75a813fbb067f4907611852c6eb1a1c1308b0df9ba5919

Observation d5957897-8f7f-41ed-977c-ebc1b3adf9fc · inbound

Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX cites this paper.

Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX Behaviour Suite for Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:40.629746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:40.629746Z digest=sha256:d943d5d7f5cd65578d557662138a3e4e9107757a98b90b5a631b7a69c9498af2

Observation 1e822db9-e1df-4e89-a247-15ccab53491b · inbound

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs cites this paper.

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs Behaviour Suite for Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:27:29.066216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:27:12.302109Z digest=sha256:9ec7693dcda6617ab99f253a49fce5f9484316388cebb7a683f473b6c2d147fa