Pith. sign in

Paper Citation Record · LEDGER

Dropout Q-Functions for Doubly Efficient Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2110.02034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.02034 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:10:01.582830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:22:15.583985Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2514533e-c7a8-473d-b31b-8ec9b4360ba1 · inbound

Learning to Play Piano in the Real World cites this paper.

Learning to Play Piano in the Real World Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:22:15.587507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:21:59.064574Z digest=sha256:bb9b233ed621b9f5e7ab490be5702acabf37d20446454c56b946d01f377c250a

Observation 9590b081-f6c2-42c2-b002-39d4d4229ea1 · inbound

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments cites this paper.

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:01.582830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:10:01.582830Z digest=sha256:e0dbcd95d1d9db3f396e7f1fb9254cd84f5b8a2fd4a31e9125fd4aef63fd23d0

Observation 962346cb-77aa-4e31-a6e4-c439b653c794 · inbound

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning cites this paper.

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:51.089069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:32:51.089069Z digest=sha256:9edce17f669649ca5fe24ff49195d8e0394bfa8771c6ba66cb6eb095af04ba7e

Observation a378b549-95f8-4ce0-962c-75f61d17b505 · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.631941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.631941Z digest=sha256:360cb15b7f3d6ad59c0dcb85d1067c974c7a7130c23d1344ce09536d413488ef

Observation 740cbb18-564c-415a-b479-54626d0a7b54 · inbound

Reinforcement learning entangling operations on spin qubits cites this paper.

Reinforcement learning entangling operations on spin qubits Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:26:15.717643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:26:15.717643Z digest=sha256:3fbdab0ff6ef5a4cfa71cd4cd0c5b679cd90174268e54726a0fb92454307c590

Observation 5d3bcf8a-fce9-414f-9150-1a6f27f5b1a5 · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:27.847603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:bed4db30462bd0460277b09b2bea64d7cc6ae19cd8c3427ebcba38623d3a2075

Observation a8075e65-fb29-4fc6-b12f-be8481159250 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.884593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:a19f403859e56c9ab85efdd524c2b244c4efb00e85bd7457140c8c91390c27dc

Observation 832ce788-dbce-43e4-95c0-4292be79cd47 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:08.493942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T14:54:30.895137Z digest=sha256:c4e44185a8d95c983d4e22f31f211933e5f0c3985ac12ba413a50fa34d83d88b

Observation 5a677ca3-750c-4b1a-93cd-c9adc57f5421 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:19:56.756183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T09:15:11.280343Z digest=sha256:a40c555e20ee6ec272d2627132525e2b9e7a0da71fe98116ea0e161edba285ff

Observation 1adb0229-71e1-45d3-b114-f389dcb864d4 · inbound

Debiased Model-based Representations for Sample-efficient Continuous Control cites this paper.

Debiased Model-based Representations for Sample-efficient Continuous Control Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.460862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:09:01.132424Z digest=sha256:cdfa5cd0c870fe9cadd6998d6bdcbf89f0ffef09d4cbf032bb34c3886c569003

Observation d8ce6f9a-7e3d-42fd-89cc-5cd12670f129 · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.055354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.055354Z digest=sha256:ff1c42366ba176aba2452f9b632e2726cc13b314cc85d7aa1bae3e9890b837a1

Observation 18432877-8ad2-4789-bcc2-c2397eaece94 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:46.394745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:46.394745Z digest=sha256:2197decd84a651a648eefe6fb0a4a807271f98676bf6dee26250bcbccb685095