Pith. sign in

Paper Citation Record · LEDGER

Off-Policy Deep Reinforcement Learning without Exploration

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:1812.02900.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1812.02900 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:12:25.792675Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T12:35:48.973163Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63b06953-ea16-4a6a-964a-e3a3380b921a · inbound

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog cites this paper.

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog Off-Policy Deep Reinforcement Learning without Exploration

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T12:35:48.976218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T12:32:09.940698Z digest=sha256:d7965eb705f52ef7de317464af5de9dabfce6c0ecbf063cbb2cec4a4ab49d3a7

Observation b57d7d31-f756-4b92-af20-246ddb4892aa · inbound

Making Efficient Use of Demonstrations to Solve Hard Exploration Problems cites this paper.

Making Efficient Use of Demonstrations to Solve Hard Exploration Problems Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T05:24:51.093084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T05:24:51.093084Z digest=sha256:2a17b8c023b1cdc6c233c05dcb475979acb4eecaa5c29c9b4197f469bbda1934

Observation b6b042a1-08d5-4666-8bde-5a1126682347 · inbound

Behavior Regularized Offline Reinforcement Learning cites this paper.

Behavior Regularized Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:19:21.423129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T15:19:21.367628Z digest=sha256:9adc3491ccb589921fb23820e5a27d6889a01872ea99ed56a597b83fe0c37339

Observation 22f1d582-6358-4ddc-9f69-86b106be883b · inbound

D4RL: Datasets for Deep Data-Driven Reinforcement Learning cites this paper.

D4RL: Datasets for Deep Data-Driven Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:19:17.409114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T23:19:17.322890Z digest=sha256:a8e066fffce421d022cb4ee2d00b12bafc744c8809d60e9238ca99dbfbd22f13

Observation 22bd0f83-d64b-421f-b16c-3c9cff3ce925 · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Off-Policy Deep Reinforcement Learning without Exploration

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:33:21.574462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:71cba31fdbca43cc0d57c73c56cb155db9073ceebe14fb742df947328ef39bf4

Observation cffc500b-6e73-442a-b4dd-fe35a2b93c91 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields Off-Policy Deep Reinforcement Learning without Exploration

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:52:43.880492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:e16179257e10dc1e7f1461c4924b8cf678e44f17fb57ff5f3bfdb7fa346c7190

Observation af49df21-5a83-47e8-8e05-54eb16dd9a2f · inbound

FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning cites this paper.

FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:25.792675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:25.792675Z digest=sha256:d14e3e5eaec8bfb6ac5fd0a6bda07c6605f5b5037354eaeffe440d0b0e8dc8c7

Observation 832ce9c9-889f-4ad2-9cb1-0298b845ddf0 · inbound

FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning cites this paper.

FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning Off-Policy Deep Reinforcement Learning without Exploration

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:49.667242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:49.667242Z digest=sha256:db52bf5254868f5d9b28ff22e0824dcfd69f2adaaebfbb614a02bbdc84c55769

Observation 838b30cd-6327-4b5e-b370-03f69d8b8d73 · inbound

Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity cites this paper.

Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity Off-Policy Deep Reinforcement Learning without Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:14:53.499845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:14:53.499845Z digest=sha256:cc0dcc52be7b06ad55e58ffdb025e7ae0c3654ab6904dfd468c4923cb59e6aa0

Observation a01e27ca-ae96-4cf7-8ea7-ad628b92dfc6 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Off-Policy Deep Reinforcement Learning without Exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.755855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.755855Z digest=sha256:d52c65c671e9e055397e0e9bb66da07a9e3840c4fb2d2bc212034cc0db946dfb

Observation d72888b8-ee2f-42b6-bb8d-62f8603480b2 · inbound

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions cites this paper.

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Off-Policy Deep Reinforcement Learning without Exploration

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:56:26.092270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T13:55:16.513839Z digest=sha256:55e4192c10536d39172bbb82e81bf08cf35915169c92f36f24f472ae4c99a7e3

Observation 59543e8c-047a-4967-8190-805e0c333707 · inbound

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management cites this paper.

Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management Off-Policy Deep Reinforcement Learning without Exploration

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:44:14.995302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T22:40:31.084594Z digest=sha256:fd7519bed75128159caca87e1f251ff075be72dbe0dadc02dd3acf7e886dbd11

Observation 205f6cf7-2b43-4d79-bc0c-0164ad81ea66 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.190747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:309254553c00c69562aef1b0cee0430ce9283c168919d531cd5d57c5cc7a7f50

Observation fb0b81a7-af25-4529-8e26-ec16c1db2bcb · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.798839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:97e5f6426096602ad34a7a5b5d8c181b757c732779f5120daf4f09cc17aabc88

Observation 2c6f62e2-ea4c-4def-abbd-e319c799d8e8 · inbound

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation cites this paper.

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Off-Policy Deep Reinforcement Learning without Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T13:15:11.461088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:15:11.461088Z digest=sha256:eacc1679bba08fc19c04ab4d7e16591179e05951352211c63e3d9e46cbc9ab9b

Observation 6416329b-4cb6-43ff-b08c-718c17b387b5 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:22.520577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:22.520577Z digest=sha256:1f238a818c6f7b10f1dbf631cfb5bc3c18b0e5a385d22de7e2d9260675400a20

Observation db5f9166-ea8d-42c8-b127-c020905a6e98 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.878960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.878960Z digest=sha256:e11d0bfe9735ee1592e006c41fa93fd9fe87ab41efc9574c9b1123a10bda427b

Observation e20d9cb9-f75b-4cdc-9654-9a97e46cb8fc · inbound

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search cites this paper.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Off-Policy Deep Reinforcement Learning without Exploration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T01:24:20.486289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:24:20.486289Z digest=sha256:21bff664690cccbcf92dddac680a2547de8e96382760a312feae561d79f22de4

Observation 9d55c0f0-601a-4975-94d9-2cc4805c1d1b · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Off-Policy Deep Reinforcement Learning without Exploration

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:30.241728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:30.241728Z digest=sha256:4756ed4ef473b1d0577f74dc160d08029249a39f76a49124ea1ed3a37bb966f2