Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2101.05982.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:15.758249Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T00:05:09.694789Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2bc50aab-90da-403e-8d65-f027559c3200 · inbound
Hadamax Encoding: Elevating Performance in Model-Free Atari Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef582d84-4c4b-44e8-9edd-05a8791a980d · inbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8f962e-0ab7-4fa2-b9c5-655623c282df · inbound
Universal Value-Function Uncertainties Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee69c96-805f-45cb-82de-044df882b3db · inbound
Safe Planning and Policy Optimization via World Model Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58b6687-ec6e-4d18-b8aa-6fc75da5b71c · inbound
The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · inbound
StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1146ce5-7daf-4a84-a69c-1cb359dba087 · inbound
A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ea33fd0-e326-43fa-8320-6b28f72fc17d · inbound
EXPO: Stable Reinforcement Learning with Expressive Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2931c806-2933-4b2c-994c-89cf5963f460 · inbound
Exploring the robustness of TractOracle methods in RL-based tractography Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a31fd7d-cfa2-471d-95dc-8e5cef1c18ba · inbound
From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 001dbc78-7a2b-4cff-a309-75d678605e99 · inbound
From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98f63cc1-5242-4820-86e2-1075fe6e3200 · inbound
What Matters for Simulation to Online Reinforcement Learning on Real Robots Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c728a3-0fd5-4f24-a6d8-3c060a8719c8 · inbound
Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41a496fc-aae2-426e-9983-0253162fa8d8 · inbound
RL Token: Bootstrapping Online RL with Vision-Language-Action Models Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5199169e-efc2-4347-9902-71254aaadce7 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c6efab4-06a8-4f45-a19c-dc765a6f0c1a · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e874d28b-ad6c-4613-a732-8fe2791b8b0b · inbound
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea3cfe6a-0723-4654-b752-a326c9d11182 · inbound
Implicit Safety Alignment from Crowd Preferences Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.