Pith. sign in

Paper Citation Record · LEDGER

Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2101.05982.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2101.05982 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:15.758249Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:05:09.694789Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2bc50aab-90da-403e-8d65-f027559c3200 · inbound

Hadamax Encoding: Elevating Performance in Model-Free Atari cites this paper.

Hadamax Encoding: Elevating Performance in Model-Free Atari Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.758249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.758249Z digest=sha256:e7804f310b3950bde5d41ebfac8188106bd580d96fa08b73446fc8e5970fae17

Observation ef582d84-4c4b-44e8-9edd-05a8791a980d · inbound

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only cites this paper.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.897231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.897231Z digest=sha256:d1969f44a6a66b2714a78b19c444c563ede989be27064f2bd07c320d17e59531

Observation 6d8f962e-0ab7-4fa2-b9c5-655623c282df · inbound

Universal Value-Function Uncertainties cites this paper.

Universal Value-Function Uncertainties Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:54.433515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:49:54.433515Z digest=sha256:5afd07b87dfd4aac2ec65b613d76b939b631528893ab74928b2483e5baff43df

Observation 2ee69c96-805f-45cb-82de-044df882b3db · inbound

Safe Planning and Policy Optimization via World Model Learning cites this paper.

Safe Planning and Policy Optimization via World Model Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:38:20.418907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:38:20.418907Z digest=sha256:7b35676b6ca60891e2aba9c352d6ad835656028fa0c168364d84e52d2c20df66

Observation b58b6687-ec6e-4d18-b8aa-6fc75da5b71c · inbound

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning cites this paper.

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:50.412668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:32:50.412668Z digest=sha256:e81ba092603bb8e06b35616dd46413db825fc556caf4e819a1ba9fbd9a2e14d6

Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · inbound

StaQ it! Growing neural networks for Policy Mirror Descent cites this paper.

StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.863748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.863748Z digest=sha256:84d013ad99d595321c36b20dfdb691a62022fdab9090a5d4c1a6f0cd0b4de595

Observation a1146ce5-7daf-4a84-a69c-1cb359dba087 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.736250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.736250Z digest=sha256:491580082b072c9742250c88aa3ce79c3225f8d35583c504e5215909229accdc

Observation 1ea33fd0-e326-43fa-8320-6b28f72fc17d · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.288471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:d97f01c828fe73641c278be8c0451bec1efe908df3dfba098c2f4824f6a3df55

Observation 2931c806-2933-4b2c-994c-89cf5963f460 · inbound

Exploring the robustness of TractOracle methods in RL-based tractography cites this paper.

Exploring the robustness of TractOracle methods in RL-based tractography Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:56.033564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:13:56.033564Z digest=sha256:4a04099d50f60a4bc4e92134548971e20ca64536e6147c820576357dcabb082f

Observation 3a31fd7d-cfa2-471d-95dc-8e5cef1c18ba · inbound

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning cites this paper.

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:34.101065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T00:36:32.681470Z digest=sha256:e30d74ca0531c66082a9a4917453c217e57eb45df9ea2940a3a92c7d52316951

Observation 001dbc78-7a2b-4cff-a309-75d678605e99 · inbound

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning cites this paper.

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:04:19.084955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T19:02:38.240098Z digest=sha256:adb12abf1b61264ce52cb217305a2bc05f3f049da838a8632adf404841e0ff6c

Observation 98f63cc1-5242-4820-86e2-1075fe6e3200 · inbound

What Matters for Simulation to Online Reinforcement Learning on Real Robots cites this paper.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.350061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.350061Z digest=sha256:5da969ae7b95a004454526ae862e6754baa7e75442673834124dff0d5d506d67

Observation 87c728a3-0fd5-4f24-a6d8-3c060a8719c8 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.870973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:5da19f612092e05e79998afd2c8c21c93a4063908b35e45d41bed83263e90e64

Observation 41a496fc-aae2-426e-9983-0253162fa8d8 · inbound

RL Token: Bootstrapping Online RL with Vision-Language-Action Models cites this paper.

RL Token: Bootstrapping Online RL with Vision-Language-Action Models Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:09.203649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T11:56:34.978806Z digest=sha256:6dd76d6f13d69973707c7b5448d4b4b6d91890eeba1f00c862a97879f009078a

Observation 5199169e-efc2-4347-9902-71254aaadce7 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:42.429745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:ed9d5fb7cdec7852c408d66b429d45cd9a746157b4f995bc63da72742b970449

Observation 2c6efab4-06a8-4f45-a19c-dc765a6f0c1a · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.696207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:d02fc2ccb4f679a61bbc44c4d333b55facb5630365d7b06501344893e802be9f

Observation e874d28b-ad6c-4613-a732-8fe2791b8b0b · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:08:59.689210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:ca5a42984181a256206a8ad9cfd4956fa64300af3c4670a295d645f3c6843bdd

Observation ea3cfe6a-0723-4654-b752-a326c9d11182 · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:31:16.890011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:3467fe3aa1ac26e870c16f2633db913818bbc51f7ae504097bc50b162ba36543