Pith. sign in

Paper Citation Record · LEDGER

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.12366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12366 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:34.992032Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff819525-73b8-41fa-b0f1-209cf67f3208 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.850630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.093584Z digest=sha256:05056fdc725ec1439a52c40ebe0938ed315808cd4a60e92b1acd8415d16a3398

Observation 62341f45-5ae2-4e76-97d0-b1477f3800f1 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.154377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.154377Z digest=sha256:b994d4f8c154af07d7573f1a5665402cbadbd77c97094705e9db5c5d95820d01

Observation d9c691f7-9c49-429c-aa7c-8f33f9ce667d · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.717508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.238890Z digest=sha256:9521c397ac06c5aa36f7850de78647ccfe66f91ce4dc9123a65516f108674726

Observation bc84f432-7f24-4c6c-bdac-2375dc1ef1b7 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.317628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.317628Z digest=sha256:f6990724ab4b523a0f192e1ef557bdde4b987db7c14ae5f659be54a6e8e5297c

Observation 62b23a52-8c10-4b28-8f78-7cf8a2647170 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Solving Rubik's Cube with a Robot Hand

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.499519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.499519Z digest=sha256:40325085e65f026a90d2292b29c0b90ddd7bcfd2f8db9f3d35512ade76c40f5b

Observation 3628204b-3c6f-4f7e-aed3-8e1bc22dc60a · outbound

This paper cites Assessing the Impact of Distribution Shift on Reinforcement Learning Performance.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Assessing the Impact of Distribution Shift on Reinforcement Learning Performance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.540733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.540733Z digest=sha256:ed3931a18cff56281563d99ab7f2fae97895b8fd111400c74376bcb94e639d78

Observation d8067b51-5833-42ea-bdd8-8ddf3eaa0c66 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.560428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.636171Z digest=sha256:c95b91a5b67b247527268fc87756bc527538379d08cb866e37aa2383bad10abc

Observation 2f0afb74-432c-4faa-af13-36455bc8470c · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.681317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.681317Z digest=sha256:9fa72e5375bf9ecf5c0b494e760af771c549f5cb11008752c48c0d200c3c688f

Observation eae48b2e-276e-42bd-a27d-a3ad041d95a6 · outbound

This paper cites & Krueger, D.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Krueger, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:37.422518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.765659Z digest=sha256:5e786a6c6270c327e8c0fbc23b9094e584582c54bb1e1329fa7d9a07ed20d1a9

Observation b85e1bf7-f3de-45f9-ab9b-376693f31989 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.229349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.800807Z digest=sha256:b4fa381c4c91c04de47866f97add38ad286548fbe10e6696f5b4dc4e173a904f

Observation f4514670-e34f-4f9d-8d03-cc0053e9feb2 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.062856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.812768Z digest=sha256:2459ed264a82c8e808e58b4fdaa28d0f437231088cd4c5e37aeccddce2ba49c3

Observation 1b0bfa77-eac1-43fc-8dea-8a54105fb0c4 · outbound

This paper cites Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:35.277560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.844964Z digest=sha256:521404510999e167a033444c7993a7b60f730579f99ef01398cc75356d07eb4f

Observation cae01744-3e5b-4045-b837-214e3d193530 · outbound

This paper cites & Fal- cone, F.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Fal- cone, F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.900260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:33.959059Z digest=sha256:37d71106236fbcaac2676fac0959cc7149c0c0de827c31c149ac68fd8eb64330

Observation 4a8df109-521d-4be5-b39c-fc5d20ccc265 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:36.705638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.070052Z digest=sha256:93dbd65213ec9fa9b3a82f66c0c90be8127e7429f9a6913bcb61167ac3afba82

Observation 05b5c852-5616-4152-b1c2-c744954b8c4e · outbound

This paper cites ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition (2024).

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.547757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.202429Z digest=sha256:f28f69be632149e640dc33c4f8abed92c916b391ac3a60249ea3ac73dc28ba94

Observation 24ff9008-bd64-4e78-9ffa-eaa4f21ff008 · outbound

This paper cites Simulated augmented reality and virtual immersive technologies.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Simulated augmented reality and virtual immersive technologies

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.419593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.334632Z digest=sha256:607090d8e000f82876c13a062912269b74554e1d053cddb3c4011188a45b19cc

Observation 2acbf66a-3a1c-4421-8c16-87365c2f07d0 · outbound

This paper cites & Yablon, Z.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Yablon, Z

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.257621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.464284Z digest=sha256:ff43ce5d94f1025cb3d3c07348a41b0ba59ee2e0f20e3fdb81470e79c674ba38

Observation 60cc0211-e0a1-4795-80a4-b7c81e549c6c · outbound

This paper cites & Luo, X.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Luo, X

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.097045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.633125Z digest=sha256:dc061bd2ecb25b79df5002df213d119d9026fac9b8dd2fc13b33d4565a3e609d

Observation c28be68d-25d9-4e01-bd62-8f058b35881b · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:35.885780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.746730Z digest=sha256:394c46365373909f0ec2d9835b945c95c03ca98cab7edaf7a5b7b2f6d008431e

Observation cb1e6732-4b91-4ebf-a565-2c1f2eb8f6a3 · outbound

This paper cites & Lee, J.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Lee, J

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:35.649468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:54:34.847455Z digest=sha256:de2da10ffe6f80fdbb6561ef1e0085ab91790b986bdf459957baf0a092d9610f

Observation 6735d5a7-b288-4389-a2a1-463332410c9d · outbound

This paper cites Unity: A General Platform for Intelligent Agents.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unity: A General Platform for Intelligent Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:34.992032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:34.992032Z digest=sha256:5838cf93863aa8d1367437a3a7de43df3aeeffc157b0d8593ea78134aff6ac83

Pith citing papers

No inbound Pith citation observations are available.