Pith. sign in

Paper Citation Record · LEDGER

Dropout Q-Functions for Doubly Efficient Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2110.02034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.02034 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:10:01.582830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:22:15.583985Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2514533e-c7a8-473d-b31b-8ec9b4360ba1 · inbound

Learning to Play Piano in the Real World cites this paper.

Learning to Play Piano in the Real World Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:22:15.587507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:21:59.064574Z digest=sha256:dbdaa2dd1fc4e00b38e8774abd66cccff294f08f80da5608b2fef2e4be0c0907

Observation 9590b081-f6c2-42c2-b002-39d4d4229ea1 · inbound

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments cites this paper.

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:01.582830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:10:01.582830Z digest=sha256:19dcb86ff6b8029c0c49ca43edfd1bb8e44ae7db934cd938efcbe4721767fcff

Observation 962346cb-77aa-4e31-a6e4-c439b653c794 · inbound

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning cites this paper.

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:51.089069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:32:51.089069Z digest=sha256:0f7fa87d5f7746d19c233ce3ffb7203181ec6f0433728189b7e0f205dc8746f8

Observation a378b549-95f8-4ce0-962c-75f61d17b505 · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.631941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.631941Z digest=sha256:9320ce6c65b72d2233fbf4c67b0439144817703e5ad7a11f28c8b7e309f3f543

Observation 740cbb18-564c-415a-b479-54626d0a7b54 · inbound

Reinforcement learning entangling operations on spin qubits cites this paper.

Reinforcement learning entangling operations on spin qubits Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T18:26:15.717643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:26:15.717643Z digest=sha256:cd995c693e0ccbebe9d88361c4bf14eeb7c6bc049464be0b16e9fcf3dfb9809b

Observation 5d3bcf8a-fce9-414f-9150-1a6f27f5b1a5 · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:27.847603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:9d7f2ed55a267b5f87f21c18eac79a72127419b9d4f5a6a4bdd26f0ad0d99e02

Observation a8075e65-fb29-4fc6-b12f-be8481159250 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.884593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:a9868b42c796889ecf6fe61e4efd71f26b735ce76c37b68d8675d7f6e244c82c

Observation 832ce788-dbce-43e4-95c0-4292be79cd47 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:08.493942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:54:30.895137Z digest=sha256:c5d7e79d0fb8979e95c48b86daf342169fee4e82589ece5c95c6d3279bf079d3

Observation 5a677ca3-750c-4b1a-93cd-c9adc57f5421 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:19:56.756183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T09:15:11.280343Z digest=sha256:4af963f0bb0db4ddd618bb09821910026eb7b393529059f9c376c2108c781ce1

Observation 1adb0229-71e1-45d3-b114-f389dcb864d4 · inbound

Debiased Model-based Representations for Sample-efficient Continuous Control cites this paper.

Debiased Model-based Representations for Sample-efficient Continuous Control Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.460862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:09:01.132424Z digest=sha256:afa2a39ef7ad245c02debe8f8420959cbe4b9fe9228d5836798bfa945c8ce990

Observation d8ce6f9a-7e3d-42fd-89cc-5cd12670f129 · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:02.055354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:02.055354Z digest=sha256:dbf533c90feed746c6424b6f633643223be2ad2a1653046031bf92cb1b501b0f

Observation 18432877-8ad2-4789-bcc2-c2397eaece94 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Dropout Q-Functions for Doubly Efficient Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:46.394745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:46.394745Z digest=sha256:530c1eefae4f5855658e0d81710240ef4c7166dc7070b35e6e67f2fe8f75c3c6