Pith. sign in

Paper Citation Record · LEDGER

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.12366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12366 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:54:34.992032Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff819525-73b8-41fa-b0f1-209cf67f3208 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.850630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.093584Z digest=sha256:f118559c8bf8e68494dbc1faeb68f91f481e918f5bcaa4bdf65dab496932f54a

Observation 62341f45-5ae2-4e76-97d0-b1477f3800f1 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.154377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.154377Z digest=sha256:6db612037441020419c5a7fba84550055c1c44111aca62351858f6c47a41f28b

Observation d9c691f7-9c49-429c-aa7c-8f33f9ce667d · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.717508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.238890Z digest=sha256:bbeaa851a5f490a46b973e0912049a4658b2e7286d5c193e4784aacf6e962e88

Observation bc84f432-7f24-4c6c-bdac-2375dc1ef1b7 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.317628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.317628Z digest=sha256:2ffa5e66eb7de839fb222123332cb1f43a7f5cd9b8ea1bf98c96a81f8d8b5132

Observation 62b23a52-8c10-4b28-8f78-7cf8a2647170 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Solving Rubik's Cube with a Robot Hand

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.499519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.499519Z digest=sha256:f251999aa0bc0dce8f9d15f870e961cfe1925cf21e0d3ead5a50d4ddb6fcf88a

Observation 3628204b-3c6f-4f7e-aed3-8e1bc22dc60a · outbound

This paper cites Assessing the Impact of Distribution Shift on Reinforcement Learning Performance.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Assessing the Impact of Distribution Shift on Reinforcement Learning Performance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.540733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.540733Z digest=sha256:01c4e78cf9fe519a5f0f5d2c11267bb6af1b7c5d881e79a83344a60caeb98d15

Observation d8067b51-5833-42ea-bdd8-8ddf3eaa0c66 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.560428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.636171Z digest=sha256:520af73f45473e7b0df75fc479c804fc9f23229549e5b2151e96d0b89e2fde5c

Observation 2f0afb74-432c-4faa-af13-36455bc8470c · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:33.681317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:33.681317Z digest=sha256:93f2071faa6ca78bc6b4f679e7489efb93c70ce79f91d92543b4f8a4250dcf8c

Observation eae48b2e-276e-42bd-a27d-a3ad041d95a6 · outbound

This paper cites & Krueger, D.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Krueger, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:37.422518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.765659Z digest=sha256:d9bd2d77f580fc4fc02a4b3792d5a3246f82de033ee10b88cf6ba721c0f6eea8

Observation b85e1bf7-f3de-45f9-ab9b-376693f31989 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.229349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.800807Z digest=sha256:5839f6d66338154abc78b4a6627d881bd822d262caeb6a87860fe6442ce23104

Observation f4514670-e34f-4f9d-8d03-cc0053e9feb2 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:37.062856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.812768Z digest=sha256:2c666651f9bee4f90776cfaa5a4f790dd2a50eb3a440b08a58b6669c8418521f

Observation 1b0bfa77-eac1-43fc-8dea-8a54105fb0c4 · outbound

This paper cites Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:54:35.277560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.844964Z digest=sha256:a2c0a33916ab4cde1ba0fe7e1234960613792ddbb645f2fc21f0e1dd8e8ebdfc

Observation cae01744-3e5b-4045-b837-214e3d193530 · outbound

This paper cites & Fal- cone, F.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Fal- cone, F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.900260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:33.959059Z digest=sha256:0cb1dd4da347dbc8f11adb611f72b2e5d5ced7a4d18f9da0d346f4a47832b413

Observation 4a8df109-521d-4be5-b39c-fc5d20ccc265 · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:36.705638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.070052Z digest=sha256:86be7158cf2d9effb40305b82acf407495829841f3e62f44d95189dd8918038d

Observation 05b5c852-5616-4152-b1c2-c744954b8c4e · outbound

This paper cites ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition (2024).

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning ARMADA: Augmented Reality for Robot Manipulation and Robot-Free Data Acquisition (2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.547757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.202429Z digest=sha256:cb72350a034eb570ce15a28f31789e43f7d0565d57e415293803f309f6107d26

Observation 24ff9008-bd64-4e78-9ffa-eaa4f21ff008 · outbound

This paper cites Simulated augmented reality and virtual immersive technologies.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Simulated augmented reality and virtual immersive technologies

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.419593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.334632Z digest=sha256:53b4f2c543b4fd672f3163d74f1ff4e602300a6c587a6d67f8b6330a32d1ea26

Observation 2acbf66a-3a1c-4421-8c16-87365c2f07d0 · outbound

This paper cites & Yablon, Z.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Yablon, Z

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.257621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.464284Z digest=sha256:ff177627af170c87ba78438076687ae472822aa98b06763e1cc16b3452691f04

Observation 60cc0211-e0a1-4795-80a4-b7c81e549c6c · outbound

This paper cites & Luo, X.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Luo, X

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:36.097045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.633125Z digest=sha256:d37c51f6d429035b24cd4919cc82ae786cd4cc26b53f51a03169ec20928635bc

Observation c28be68d-25d9-4e01-bd62-8f058b35881b · outbound

This paper cites an unresolved cited work.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:54:35.885780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.746730Z digest=sha256:8d69e077b9dd3ff604c2b56c82f31feb6354c2edbc86669f4a513c42e2f4365a

Observation cb1e6732-4b91-4ebf-a565-2c1f2eb8f6a3 · outbound

This paper cites & Lee, J.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning & Lee, J

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:54:35.649468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:54:34.847455Z digest=sha256:b0914bccb5fabc3bab61eeac791adfc339465905919ef5369a0f04d5a70a8494

Observation 6735d5a7-b288-4389-a2a1-463332410c9d · outbound

This paper cites Unity: A General Platform for Intelligent Agents.

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning Unity: A General Platform for Intelligent Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:34.992032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:34.992032Z digest=sha256:afbe76c165e7896a7949fe0f45e704683f1a9eb3ff91194405f193d423964938

Pith citing papers

No inbound Pith citation observations are available.