Pith. sign in

Paper Citation Record · LEDGER

Where to Intervene: Action Selection in Deep Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.04187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04187 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:58:21.440862Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact4
  • verified fuzzy8
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9df342b2-8219-4261-af34-b8175019f888 · outbound

This paper cites C.2 Treatment Allocation for Sepsis Patients We utilize the MIMIC-III Clinical Database to construct our environment for Sepsis patients.

Where to Intervene: Action Selection in Deep Reinforcement Learning C.2 Treatment Allocation for Sepsis Patients We utilize the MIMIC-III Clinical Database to construct our environment for Sepsis patients

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:23.609894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:21.134632Z digest=sha256:b7143978875b45b6a5d56e1d3b92d06ac6dad3fe39cb9e74cff0c3357b8dd32d

Observation e8bf660f-5299-4e58-8a57-fafc58285a48 · outbound

This paper cites This condition is typically met by standard tabular machine learning algorithms.

Where to Intervene: Action Selection in Deep Reinforcement Learning This condition is typically met by standard tabular machine learning algorithms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:23.131846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:21.335508Z digest=sha256:116411dbbb84fc81f8038da082a37eda947c506bed572f648a4540b3f927f710

Observation ec79a4de-c4d4-480d-a0eb-a4465a56a4b2 · outbound

This paper cites Model-Based Reinforcement Learning for Atari.

Where to Intervene: Action Selection in Deep Reinforcement Learning Model-Based Reinforcement Learning for Atari

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:19.733463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:19.733463Z digest=sha256:47a8391a9d55aed0f76da75a2c68988ef5ad8afa15383c2d6a4bd50123825848

Observation f9266fcf-3437-40ca-a0a5-1007293cd133 · outbound

This paper cites Quasi-optimal Reinforcement Learning with Continuous Actions.

Where to Intervene: Action Selection in Deep Reinforcement Learning Quasi-optimal Reinforcement Learning with Continuous Actions

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:58:22.385863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:20.217655Z digest=sha256:3a5886675350e5fd6a27b0bb5e79407fffde360c3e51ae59e4b672720bfd3c46

Observation f7c6235a-514c-4b58-9092-c03351d87ec6 · outbound

This paper cites Sequential Knockoffs for Variable Selection in Reinforcement Learning.

Where to Intervene: Action Selection in Deep Reinforcement Learning Sequential Knockoffs for Variable Selection in Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:58:22.102178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:20.447001Z digest=sha256:d390ec26a7d9a2387b3a4fd7ec20356a1704bea6392da8b63af7d1895fd1d5b1

Observation 2b4f1771-5967-4d9a-8cd9-0fd7c74dfb36 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Where to Intervene: Action Selection in Deep Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:20.695847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:20.695847Z digest=sha256:e8e2a6ac71361a5361556c5e91042f31be0f86dd120558f137fedfc398b3c4c8

Observation 71d44aa7-5877-4d8c-92cb-277a1e0e7ac9 · outbound

This paper cites FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback.

Where to Intervene: Action Selection in Deep Reinforcement Learning FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:58:21.676126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:20.970823Z digest=sha256:d606e90e300f9a8e290ff7058de320b92cf0c53736e55136c08166eaee548202

Observation afd1b8d0-f784-4674-ad9a-4e02b89a5410 · outbound

This paper cites an unresolved cited work.

Where to Intervene: Action Selection in Deep Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:58:23.384146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:21.231112Z digest=sha256:4ee4f3a9f34d1288e2c251862bed22166a9074cf3e47348b4b6d82b8ad3999d5

Observation 790f878e-7e0c-4988-9cee-407c0419c501 · outbound

This paper cites Then for suchϵ, denote Ω :={i :ϵi =−1}, which is a subset ofH0 by the assumption (and recall thatH0 is the collection of all null variables).

Where to Intervene: Action Selection in Deep Reinforcement Learning Then for suchϵ, denote Ω :={i :ϵi =−1}, which is a subset ofH0 by the assumption (and recall thatH0 is the collection of all null variables)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:22.869856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:21.440862Z digest=sha256:115460e47e4bc602cd61447f58ca03de9d8a7b3e72b3f6def97481b6a31bf670

Observation 9378e9b2-1a72-447f-8335-7fa9a1e3eff7 · outbound

This paper cites Generalized Fisher Score for Feature Selection.

Where to Intervene: Action Selection in Deep Reinforcement Learning Generalized Fisher Score for Feature Selection

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:19.447628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:19.447628Z digest=sha256:a49e0a61718f6ba69ce29b5dced5abcc67b60d5d65b670adc5fd4937d5397ea6

Observation 31e1abca-1c61-4ab8-83b0-a0716e62a01e · outbound

This paper cites Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions.

Where to Intervene: Action Selection in Deep Reinforcement Learning Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:20.864765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:20.864765Z digest=sha256:707cd71f6fb1c691261c4dd0830dba63a372d5f0a0e6cdbb7fe389a779ace429

Observation b809d0b4-1ece-45db-b4af-c2a9d1f658af · outbound

This paper cites Sample Efficient Feature Selection for Factored MDPs.

Where to Intervene: Action Selection in Deep Reinforcement Learning Sample Efficient Feature Selection for Factored MDPs

Reference 2012

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:58:22.570256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:19.576991Z digest=sha256:f125615b215018d06c20431097f2df4bca903db269549a5459e9c2ab178e4a90

Observation b228eede-2d25-4c27-b197-7cad3bac34da · outbound

This paper cites Modern perspectives on reinforcement learning in finance.Modern Perspectiveson ReinforcementLearning in Finance (September 6, 2019).

Where to Intervene: Action Selection in Deep Reinforcement Learning Modern perspectives on reinforcement learning in finance.Modern Perspectiveson ReinforcementLearning in Finance (September 6, 2019)

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:24.254348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:19.989486Z digest=sha256:c8de1aee75334262322ab3987c24ea040458137196413d41aff9538b5f904800

Observation 16e60d3e-f86f-49f5-b7a5-fbdd0d0b02b8 · outbound

This paper cites Growing action spaces.

Where to Intervene: Action Selection in Deep Reinforcement Learning Growing action spaces

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:24.587546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:19.344036Z digest=sha256:17161c513008fa7253c728f67c62a1992b5c51adf4db7f8bbc58543a97ef0f29

Observation 56d90c08-b572-469f-96eb-ba29dc9da6cd · outbound

This paper cites Action space shaping in deep reinforcement learning.

Where to Intervene: Action Selection in Deep Reinforcement Learning Action space shaping in deep reinforcement learning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:24.408502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:19.856059Z digest=sha256:dc1c8542b31431eb7ccb01acc3ce9d8090a1c89b83704871100bb59c77bfd84c

Observation 94004926-3825-4132-9d68-ab21500bb345 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Where to Intervene: Action Selection in Deep Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:20.548361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:20.548361Z digest=sha256:7ad1cc117dde594f20a38808ca21406dc51c3cd3124f7a17dacf4e1d907ba15d

Observation 602f891a-aa8c-49bb-874c-e95ad1ecf200 · outbound

This paper cites Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling.

Where to Intervene: Action Selection in Deep Reinforcement Learning Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:24.078551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:20.102586Z digest=sha256:a35c76e5e90ffec19aa4fda56561da74d3b19d74b72462b81a14d599668d2880

Observation 072990f4-7d05-4765-aeb1-0fd3c5c074a4 · outbound

This paper cites Auto-Encoding Knockoff Generator for FDR Controlled Variable Selection.

Where to Intervene: Action Selection in Deep Reinforcement Learning Auto-Encoding Knockoff Generator for FDR Controlled Variable Selection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:20.357047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:58:20.357047Z digest=sha256:80ddeafd7927f21fef9d47600e06180dc04ef6ce7a92dec287773abf719b3f73

Observation caf74bf3-4425-4c58-8be6-17dff894d6c5 · outbound

This paper cites Gene Hunting with Knockoffs for Hidden Markov Models.

Where to Intervene: Action Selection in Deep Reinforcement Learning Gene Hunting with Knockoffs for Hidden Markov Models

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:58:21.857131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:20.777883Z digest=sha256:0c07607d133d6cc210295b0737c27a7cbb96c7c90c59624c188475167ff1474a

Observation f5db8414-e75c-4636-956e-1f35771c0f4f · outbound

This paper cites (2023) adopted a two-stage framework, performing variable selection offline before applying reinforcement learning.

Where to Intervene: Action Selection in Deep Reinforcement Learning (2023) adopted a two-stage framework, performing variable selection offline before applying reinforcement learning

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:58:23.886205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:58:21.040154Z digest=sha256:1e79750bb3f08e735aafcf0db705925ee11f2b3be32d67bf1f3741d5fb5ad0de

Pith citing papers

No inbound Pith citation observations are available.