Pith. sign in

Paper Citation Record · LEDGER

A Survey on Explainable Deep Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 6 inbound Pith citation observations for arXiv:2502.06869.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06869 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:17:22.041681Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:07:28.048575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:42.446579Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e81ce1a-1e00-4754-b65a-55f637c05cf9 · outbound

This paper cites CDT: Cascading Decision Trees for Explainable Reinforcement Learning.

A Survey on Explainable Deep Reinforcement Learning CDT: Cascading Decision Trees for Explainable Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.993559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.993559Z digest=sha256:9deccd291e2485ea4b32066aac24b9871853ccb7685afa36704d68eda74b8b9a

Observation ee3e32a0-d58f-45bc-a7e4-2ae32c741b4f · outbound

This paper cites Empirical influence functions to understand the logic of fine-tuning.

A Survey on Explainable Deep Reinforcement Learning Empirical influence functions to understand the logic of fine-tuning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:17:22.150370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:22.016630Z digest=sha256:436e4a5b33d66aab447291fe31655b5e046a2c7868fe9e90805defb968100d8a

Observation ec979f0d-0b29-46bb-9a79-ab16384dd96d · outbound

This paper cites Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models.

A Survey on Explainable Deep Reinforcement Learning Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.027145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.027145Z digest=sha256:c988ffcdb765acdd69418c691569d20b4627ff90dbc3aaffc7b167e54923f3c6

Observation ad3a76e1-43dc-43d6-939d-1872bc260bf3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Survey on Explainable Deep Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.032072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.032072Z digest=sha256:4bc5744165aeced642ec70945383e19bcb6d3d282c6b3acbb5c53b1180e9b73a

Observation db1f2503-fc15-4377-983f-8d34455fe331 · outbound

This paper cites StarCraft II: A New Challenge for Reinforcement Learning.

A Survey on Explainable Deep Reinforcement Learning StarCraft II: A New Challenge for Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.036708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.036708Z digest=sha256:200aaf77b71567f5b625ab9380da967c6ec4ce539394e4cc81e0dc7458eda9c4

Observation 20ad6279-a0ce-4010-b33c-c23b7aeb18d5 · outbound

This paper cites Graying the black box: Understanding dqns.

A Survey on Explainable Deep Reinforcement Learning Graying the black box: Understanding dqns

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:17:22.261860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:17:22.041681Z digest=sha256:a442b6680f6bd43a3d28cc635669687cb98118e643de26ad281d3c5a43358552

Observation bf3fe872-86b8-4ab3-85ea-41bf7df95140 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

A Survey on Explainable Deep Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.022282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.022282Z digest=sha256:bd88a3998ebd9c8c5f95f794b81c0ecb7feb08790991b3d7b5c781356783f40b

Observation b8418f0f-c047-4149-90ec-39cef4f31e85 · outbound

This paper cites Do Influence Functions Work on Large Language Models?.

A Survey on Explainable Deep Reinforcement Learning Do Influence Functions Work on Large Language Models?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.010430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.010430Z digest=sha256:059c2e31a9cb0374d9e73ae9cc4c222dc893967267ceabd6d6e86ab8ff2a09cd

Observation a07730d5-54b2-44b5-9f9f-6985f2483595 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

A Survey on Explainable Deep Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.982058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.982058Z digest=sha256:e12e4961454ce5a6f0f7a7768357971a45ee1bbeb5e1eb7398ca3f898457c52b

Observation 639bc1e6-02bb-4c76-aff6-82d189ed02b8 · outbound

This paper cites Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards.

A Survey on Explainable Deep Reinforcement Learning Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:22.005131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:22.005131Z digest=sha256:3ad8a6d9554a5e042c98ad14a77a7502041b274e15a76ed76dadd70c0f80dedb

Observation 8e5c1f63-3ef3-46c3-a3c9-2ffc12ecd55e · outbound

This paper cites PromptExp: Multi-granularity Prompt Explanation of Large Language Models.

A Survey on Explainable Deep Reinforcement Learning PromptExp: Multi-granularity Prompt Explanation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.999381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.999381Z digest=sha256:78edcf918eee8dd380867e214e654eb6f077669dfd129118fe73c0c604fb97fd

Observation c662d386-ef9e-417a-a2b9-e734c20dced9 · outbound

This paper cites Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models.

A Survey on Explainable Deep Reinforcement Learning Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T19:17:21.988056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:17:21.988056Z digest=sha256:3084313243d22a5c28c1d21b75b42bebef77ea8e11f9ed63ceb90dc0766b34b4

Pith citing papers

Observation c0ba9a2d-127b-4f21-8246-65e015834008 · inbound

Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration cites this paper.

Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration A Survey on Explainable Deep Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:28.048575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:28.048575Z digest=sha256:3a9690ff1fd6cd775e6280b9c85aae19e0cd76dea51c078191f8eb673b787054

Observation f13abb6d-7170-4bb9-be00-4edc0e906c68 · inbound

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI cites this paper.

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI A Survey on Explainable Deep Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:08:49.409444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:08:49.409444Z digest=sha256:8847d4e8d3de5fe896d6c7d4d3c92b33f96de1519641209bb45b46034e7a4d95

Observation 0712794e-c8e0-494b-bdb3-83cfa4fc9922 · inbound

Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments cites this paper.

Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments A Survey on Explainable Deep Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:46:26.651930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:46:26.651930Z digest=sha256:f1cde930621ef4ab8fff47a891ce45617d1f5f629f3971a522ff4ae02313b6ec

Observation 59084918-d9da-4c68-8672-3270b04e54bd · inbound

M2-PALE: A Framework for Explaining Multi-Agent MCTS--Minimax Hybrids via Process Mining and LLMs cites this paper.

M2-PALE: A Framework for Explaining Multi-Agent MCTS--Minimax Hybrids via Process Mining and LLMs A Survey on Explainable Deep Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:50:20.792149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:42:03.298417Z digest=sha256:97d7225e4353221570c332bcac7b7a6de6976e555174363b92d18173c39ccb79

Observation 537885f0-3e60-47b4-814f-40b0895d394c · inbound

Price of Fairness in Short-Term and Long-Term Algorithmic Selections cites this paper.

Price of Fairness in Short-Term and Long-Term Algorithmic Selections A Survey on Explainable Deep Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:11.156182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T10:04:04.849162Z digest=sha256:cd3a038c029679b65a351f4faa36054d1abd07fe3523802797ff025a0cce0455

Observation 05e54b9f-39a5-4713-b01b-17b1f07fc09e · inbound

A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI cites this paper.

A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI A Survey on Explainable Deep Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:42.448163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:07:50.847698Z digest=sha256:f5430333e2036610762c8b61b95b143fc68bd48d88addb7160d01cdbd4592eea