Pith. sign in

Paper Citation Record · LEDGER

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

As of 17 August 2026, this Paper Citation Record lists 100 of 106 outbound references and 1 inbound Pith citation observation for arXiv:2412.05783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05783 v1

Coverage vector

measured 100 of 106 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:26:49.464366Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:36:27.422590Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 106 outbound references displayed

  • verified exact8
  • verified fuzzy32
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5cf9ca6b-303e-49b7-860a-b97f038ec39c · outbound

This paper cites Formulation and estimation of dynamic models using panel data.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Formulation and estimation of dynamic models using panel data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.013868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.013868Z digest=sha256:6c569f0fad719fe539577c15ce904e09cef710c4334283743c57a5aa96d26da9

Observation 81eab347-b38c-4953-a397-f6fa7bbcb21d · outbound

This paper cites Doubly robust identification for causal panel data models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly robust identification for causal panel data models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.019604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.019604Z digest=sha256:43e60379daa5c681399347cb123b5c101c195bf27c532d4683a11f641c8c8918

Observation d392ba94-52de-418b-9353-b0106c23bf1f · outbound

This paper cites Design-based analysis in difference-in-differences settings with staggered adoption.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Design-based analysis in difference-in-differences settings with staggered adoption

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.024472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.024472Z digest=sha256:2bcf13a306d6e326b2cb62f14cbe17f6fa5dc007a8655a44661caafdd4bd17a6

Observation a13dec1d-9588-44aa-bf28-610180a1d069 · outbound

This paper cites Econometric analysis of panel data, volume 4.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Econometric analysis of panel data, volume 4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.029217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.029217Z digest=sha256:8e915de337cf88b24165ee08c6b3db2d0d3f67800b555defda31a6bcd538fd2d

Observation 856b2df3-9142-4198-a91c-0af226994e1d · outbound

This paper cites Genetic risk profiles for cancer susceptibility and therapy response.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Genetic risk profiles for cancer susceptibility and therapy response

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.033801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.033801Z digest=sha256:7480fd73a9f173c43ec489317cfe563b4ce7d87daa2e7c86db409600325aab1e

Observation b4fc580b-0d56-486a-9196-350a637475a3 · outbound

This paper cites Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.038799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.038799Z digest=sha256:8fed7462340f78ba4dbeb6d63cb533e4bdcf40065d9dfb8b2e5255a1861f97ec

Observation 2d2851fd-f9da-422f-a698-a9b11f3f2593 · outbound

This paper cites Off-policy Evaluation in Doubly Inhomogeneous Environments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy Evaluation in Doubly Inhomogeneous Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.168208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.043818Z digest=sha256:45a4a0164ab6d3b6049a8aa927199a6185558cec6f4ee14f4fe61b1663039c09

Observation e6aa3b11-ad83-48f3-aa5e-d6c55cdf5ca9 · outbound

This paper cites Time series deconfounder: Estimating treatment effects over time in the presence of hidden confounders.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Time series deconfounder: Estimating treatment effects over time in the presence of hidden confounders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.048680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.048680Z digest=sha256:d13c4c0286b3ef9579ae00dbee96a88d6479ce5dc595e30fa87b0e999eb196cb

Observation f3703407-ced8-4883-8a64-918f469106e4 · outbound

This paper cites Robust fitted-q-evaluation and iteration under sequentially exogenous unobserved confounders.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Robust fitted-q-evaluation and iteration under sequentially exogenous unobserved confounders

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.052749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.052749Z digest=sha256:3e86cd960074c462614682df865a0477e3188311304ef58d1368f508cd10fd3c

Observation 2f731838-1fa3-45b9-bd8d-fc1dcab8f1d9 · outbound

This paper cites Treatment effects in interactive fixed effects models with a small number of time periods.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Treatment effects in interactive fixed effects models with a small number of time periods

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.057216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.057216Z digest=sha256:883995f91619e3eb5ef7da68c8cd52da1b2792839d6a639a3a2c0d338c342279

Observation 64d47dda-039d-4779-9afd-7c65d9605d74 · outbound

This paper cites Statistical inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Statistical inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.061405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.061405Z digest=sha256:279810ab8ef579b8077bd6c51cb1d17596975e6b0c04faeb0d6f1b0bb177eca6

Observation 674bde1d-3223-409c-babe-49b6fe6a5f80 · outbound

This paper cites Universal off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Universal off-policy evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.065314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.065314Z digest=sha256:e06cb85773f3a2a48df38f7e66bd078a8c2654df14a5f567ff0a52c8aa8ce423

Observation ea0ba1f1-cf76-4fa8-ad75-515556cd94b5 · outbound

This paper cites Testing for the markov property in time series.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Testing for the markov property in time series

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.069222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.069222Z digest=sha256:6d2c731a991ce794088f4f065c15bb4bee959c4fcfea6cedec98d697d3a33502

Observation 61ddd041-b55b-4d1e-932d-44105215db94 · outbound

This paper cites Information-Theoretic Considerations in Batch Reinforcement Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Information-Theoretic Considerations in Batch Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.069475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.073582Z digest=sha256:b22d7c1640d4e36fc6de4398e7368dfa9dc642253119653e2f662b131340d441

Observation 61dbe7c9-4d78-4c87-b6a3-b8ea0c5428b2 · outbound

This paper cites On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.078519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.078519Z digest=sha256:4495ceb6566245272a4e6741973acdfdcc76b3e048989e34477334c3014d4651

Observation 79a87a3e-8b07-4c7e-9a83-f545bcf34291 · outbound

This paper cites On instrumental variable regression for deep offline policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On instrumental variable regression for deep offline policy evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.083127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.083127Z digest=sha256:cecd5c9986682a290f42087e99f656c3647e0f8cdf616767b3b9f78bde05242a

Observation cca7315a-4c7e-47dc-9384-2d46fd4f03d5 · outbound

This paper cites Double/debiased machine learning for treatment and structural parameters.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Double/debiased machine learning for treatment and structural parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.087629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.087629Z digest=sha256:bcf4de174a5318e6f519a027dda773b4aba65601e40e684e85eadc84101a558c

Observation 54065e70-f661-480c-b050-cb993b19469e · outbound

This paper cites A crash course in good and bad controls.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A crash course in good and bad controls

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.092271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.092271Z digest=sha256:40224d2f0d035cae2a1f3c4511b7316f8b58d0b53339677216005e76c0fb2f58

Observation b3647527-68eb-42b6-ad3c-2c247dc2fb07 · outbound

This paper cites Coindice: Off-policy confidence interval estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Coindice: Off-policy confidence interval estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.096926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.096926Z digest=sha256:ce4297d00ffd3cbc9ab96748c8c575fb753c6324e08e12e09e2564974b5b4b5b

Observation ad05cb02-d180-4a4a-9508-f57b7f20b71e · outbound

This paper cites Comment: Reflections on the deconfounder, 2019.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Comment: Reflections on the deconfounder, 2019

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.101485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.101485Z digest=sha256:214595ab43bad22ff15f0049f5e9d37719eb04280f416bd75950ce25bbc4e63d

Observation 68dde355-1d55-49a3-b452-15c5bef507d5 · outbound

This paper cites Two-way fixed effects estimators with heterogeneous treatment effects.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Two-way fixed effects estimators with heterogeneous treatment effects

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.105859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.105859Z digest=sha256:996e2ea9630849b6e2854aa72b42cdd83c20b88ad1a40bddb188d8ac5901ad59

Observation 9ea2261c-b5ee-4ba1-aa62-e36811700efe · outbound

This paper cites Counterfactual inference in sequential experiments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Counterfactual inference in sequential experiments

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.049057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.110445Z digest=sha256:997d436c56200ee2d205704dd4529912015cebae369127dcaa6d6514541a5ee5

Observation 722ea65d-52ee-42d8-9ce1-5a3902d93c9f · outbound

This paper cites A theoretical analysis of deep q-learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A theoretical analysis of deep q-learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.115428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.115428Z digest=sha256:6f58283b15debb60fc3fcd897f5037512a0d099a9d8963f359bf7f71ce33c8ee

Observation 6dc54f2f-3ac3-4e0a-ba59-30af9c519f98 · outbound

This paper cites More robust doubly robust off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning More robust doubly robust off-policy evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.119846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.119846Z digest=sha256:ad59c66cea5e2516242190ab070ebded07be7ae828171180efd2cdc90b496183

Observation 1aa730a7-d898-46c9-9e90-bd224f8486c2 · outbound

This paper cites Deep neural networks for estimation and inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Deep neural networks for estimation and inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.124324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.124324Z digest=sha256:00e92da27fa99cf4e2396e4dc6ec1f265b1e3b5fb4b292add3b3ada01977eda8

Observation d241ef39-edf2-4287-b0e8-deb304ea1fc6 · outbound

This paper cites Non-parametric panel data models with interactive fixed effects.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Non-parametric panel data models with interactive fixed effects

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.128856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.128856Z digest=sha256:f651a4a25a3e83a89b7d6019a3eb3baddd7c9a948698e04fe11bc7ece674be78

Observation 4e1aa98c-71ab-44ca-8711-72bbcd9c208f · outbound

This paper cites Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.028564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.133163Z digest=sha256:1b88584816f9fda0cd1d4255ba931ce99b233393e3b93ce959085da627c55a1c

Observation 71f7fafc-e76d-4c41-856e-74588a2a464e · outbound

This paper cites Prediction of treatment response for combined chemo-and radiation therapy for non-small cell lung cancer patients using a bio-mathematical model.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Prediction of treatment response for combined chemo-and radiation therapy for non-small cell lung cancer patients using a bio-mathematical model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.138132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.138132Z digest=sha256:ec164de5f746f41a4b88c68fdbdcb459f23e8742dbe6b8603bf57c9637c55f39

Observation af85bcdc-9ff7-4a0b-bd0c-dc89ad0653d8 · outbound

This paper cites Issues in assessing the contribution of research and development to productivity growth.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Issues in assessing the contribution of research and development to productivity growth

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.142440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.142440Z digest=sha256:bd529bb5c154312bd8c38a3686732cda71823f5c0e08884092e77f07bc095ab1

Observation 4ca90b1d-0ee0-404d-a32e-8736af8529a9 · outbound

This paper cites Richard Guo, Anton Rask Lundborg, and Qingyuan Zhao.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Richard Guo, Anton Rask Lundborg, and Qingyuan Zhao

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.147024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.147024Z digest=sha256:f7a9af280ec8b014ea5a0946a1ade698a17f62f7334ad6cc034c59ca20dc7dc6

Observation 52c89603-207b-4f0a-b66d-ed48e4f6765f · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Dream to Control: Learning Behaviors by Latent Imagination

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.151455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.151455Z digest=sha256:6c63b447eee626e3dce9d8defb241166c1641b53b56034fee224aa922fd1c742

Observation 3bfd73e2-7b4c-4d1b-a9ba-66f424d966f0 · outbound

This paper cites Learning latent dynamics for planning from pixels.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Learning latent dynamics for planning from pixels

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.156301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.156301Z digest=sha256:83625e3d3461ef73054d814829daf709ceb19006181e95c0b2698b1cc23b8930

Observation fa09841b-afa9-4af0-8f09-c107a472fb7d · outbound

This paper cites Bootstrapping fitted q-evaluation for off-policy inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Bootstrapping fitted q-evaluation for off-policy inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.160798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.160798Z digest=sha256:06178be8fb17ea83d4fdb5fe6eb13b864454527daab2da5358f9a0fa876e72db

Observation 52f1e374-03d7-4990-9a8d-02939d46e023 · outbound

This paper cites Sequential Deconfounding for Causal Inference with Unobserved Confounders.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Sequential Deconfounding for Causal Inference with Unobserved Confounders

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:49.994127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.165490Z digest=sha256:7284034e85688400896c57c581681aa180084e5212c69947650482f570af1f31

Observation 7fe94afa-707d-4d74-8f53-c5f8b8746612 · outbound

This paper cites Neural collaborative filtering.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Neural collaborative filtering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.170520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.170520Z digest=sha256:3f8bb2624d6bd75ba0d4c5d8c8dd8d0c952bd10663698ef6cf8d73b6db974299

Observation 1b616d7c-1356-4f6b-8bf4-a49d8e0f5007 · outbound

This paper cites A Policy Gradient Method for Confounded POMDPs.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Policy Gradient Method for Confounded POMDPs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.174755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.174755Z digest=sha256:da961db02037dafe930a737fd8cdc0044c7ae6130c7cf6f8644126a3796fb845

Observation e4cadad5-e646-4614-a495-325ffc3a13c8 · outbound

This paper cites Collaborative filtering for implicit feedback datasets.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Collaborative filtering for implicit feedback datasets

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.178967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.178967Z digest=sha256:5f104acd6e1cec5b865933633de2668c308cae4b6924693b935b3bde4ad2239f

Observation f18883af-ed44-4dd2-9508-340088c4bad5 · outbound

This paper cites On the use of two-way fixed effects regression models for causal inference with panel data.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On the use of two-way fixed effects regression models for causal inference with panel data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.182754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.182754Z digest=sha256:c2cea94b26616cb41b3a24e3bdfb16c7f636eb8fdc74b7aedc8b0e5729c5073d

Observation 32386744-8488-4fe7-9487-63338c368b07 · outbound

This paper cites Off-policy evaluation via off-policy classification.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy evaluation via off-policy classification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.799584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.186790Z digest=sha256:00650bfd69a7f242cdad6b8566e924a38b1fb7382d5f3d08856a4a7492c94083

Observation a127d31e-b7f1-4795-8641-129d4520de1b · outbound

This paper cites When to trust your model: Model-based policy optimization.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning When to trust your model: Model-based policy optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.190846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.190846Z digest=sha256:e9f5e25c8550e601edb6a04bba34d51e2198a7c595f646cb918641a59eb98204

Observation 71aaaba8-9046-4a8e-94b2-8f8cf63173aa · outbound

This paper cites A survey on knowledge graphs: Representation, acquisition, and applications.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A survey on knowledge graphs: Representation, acquisition, and applications

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.194737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.194737Z digest=sha256:3b403c9ade88525b74a82b0be3d9d95bf7af6795f4d10e40fbb872d8ff8dd7b8

Observation 0b532d18-1e37-4f82-a1f3-29eeea0c962c · outbound

This paper cites A Note on Loss Functions and Error Compounding in Model-based Reinforcement Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Note on Loss Functions and Error Compounding in Model-based Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.199064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.199064Z digest=sha256:655995be506fae4f81646fa275a946614357cc7cb8f7a1690e839d75bd9976c0

Observation a49847dc-e071-44ab-8e47-3da895f4dd04 · outbound

This paper cites Doubly robust off-policy value evaluation for reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly robust off-policy value evaluation for reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.766712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.203564Z digest=sha256:641d9715d18dcc7498de7fe25347ee59f07f2080f396b21aa1f77aa0c83ef202

Observation af55c0cd-0b26-43a5-a67a-2265db6863e3 · outbound

This paper cites Mimic-iii, a freely accessible critical care database.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Mimic-iii, a freely accessible critical care database

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.208020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.208020Z digest=sha256:747f87b26cc47828620b2f49c3727d4d876d2bc77537b8bc331a900fb65cb891

Observation 64ee11ac-76dd-476a-87f0-6ae58685e7f9 · outbound

This paper cites Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.212413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.212413Z digest=sha256:5a672f6dec6335676f7bb1550f748bd9971a751b6f1005cd6ffb9fe4b6d084fa

Observation 0e964bf1-c5cb-41ef-a054-15e682554714 · outbound

This paper cites Confounding-robust policy evaluation in infinite-horizon reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Confounding-robust policy evaluation in infinite-horizon reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.734020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.216658Z digest=sha256:b5e32b28bad869b78eb6baed609ed8812f0214d79b49294cf2ac2532d9ec2f72

Observation 94a0e4f1-6eb6-483d-a136-302449c5190e · outbound

This paper cites Offline Policy Evaluation and Optimization under Confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Offline Policy Evaluation and Optimization under Confounding

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:49.943952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.221467Z digest=sha256:fce3a496b02ece85f1b508be79df81af5ca31da7b1d2897d9bc2ba309937a079

Observation 3b6da3e7-58ef-4f93-be19-d93779cdf21d · outbound

This paper cites Learning mixtures of markov chains and mdps.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Learning mixtures of markov chains and mdps

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.719425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.227412Z digest=sha256:6ebc443b80a9b9bbdb11134706b618db6971ec053aba1f6f10c2d0125218ea19

Observation 0e270572-5815-4174-bd97-59858a67f107 · outbound

This paper cites Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.232155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.232155Z digest=sha256:d85cb029f20284651157c77772ffafc750158d13ae85bff7ff39049b9fc8e3d4

Observation ff919683-2e4f-47ab-94b7-05a99ee2375c · outbound

This paper cites Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.237033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.237033Z digest=sha256:975ca1d09507b376ee445cd63a0d1c156055b76d5cdeb78c39d8622c403aba35

Observation d548c79c-5569-492f-bacf-4f08e31279a2 · outbound

This paper cites Off-policy estimation of long-term average outcomes with applications to mobile health.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy estimation of long-term average outcomes with applications to mobile health

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.705579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.241606Z digest=sha256:fcd35d54b312a0527c5190c274a8e1a2739b93e99b3df3d35b914a5488e01174

Observation b53a4e2a-c74e-4d51-9b97-02e1eeb4febb · outbound

This paper cites Batch policy learning in average reward markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Batch policy learning in average reward markov decision processes

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.690923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.245977Z digest=sha256:36c3b3e83ab0869da8b42f2895738d2e7430dc196aa535fc8030543d32135dfb

Observation 4e5d4ad9-54f4-44a1-ac39-e5d5a85ef690 · outbound

This paper cites Forecasting treatment responses over time using recurrent marginal structural networks.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Forecasting treatment responses over time using recurrent marginal structural networks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.250321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.250321Z digest=sha256:9bbd7e61a24e9aa3e318dab5cb74ce1aee8019636d44b189e2efed24d3db5ba8

Observation 65385c22-c2a2-4eff-b63e-d9162e098533 · outbound

This paper cites Breaking the curse of horizon: Infinite-horizon off-policy estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Breaking the curse of horizon: Infinite-horizon off-policy estimation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.254963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.254963Z digest=sha256:7be643786ec5f07e2a598e5ec4cef8a4f100b8d23662560879042c1f5be150a3

Observation 0cc27e4d-b8ac-4b9f-bd0f-8969dd803f30 · outbound

This paper cites Provably good batch off-policy reinforcement learning without great exploration.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Provably good batch off-policy reinforcement learning without great exploration

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.658294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.259473Z digest=sha256:462978cca2a71946ea371bd0aeff776102319b652025ff0ad6dbff288aef4196

Observation 6020fd7c-11f3-408c-8ef3-d621c217b820 · outbound

This paper cites Causal effect inference with deep latent-variable models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Causal effect inference with deep latent-variable models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.264028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.264028Z digest=sha256:52accb5dc033f5e640ea02cf1241007164ffea934f101ce9139e193717c007af

Observation 011c4092-9749-427a-9ef8-26338579198c · outbound

This paper cites Deconfounding Reinforcement Learning in Observational Settings.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Deconfounding Reinforcement Learning in Observational Settings

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.268555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.268555Z digest=sha256:9fc2c4ec35009fb45ac5d3a97911a5158d8cfdc9275f573e2d70b3141e7ef84d

Observation fd481239-bf69-4c47-9a24-db357de512c5 · outbound

This paper cites Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision processes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.636096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.273622Z digest=sha256:41b340c04e885606c605249df96b79f52c5e151d52e63dcf83843dd1775c7d9f

Observation 4a51b86d-dc62-426d-afdc-9fa3765d7c6b · outbound

This paper cites Estimating causal peer influence in homophilous social networks by inferring latent locations.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Estimating causal peer influence in homophilous social networks by inferring latent locations

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.623463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.278278Z digest=sha256:f4e427fdc237d862fbf6b225cc9c7d41c34aa79e756649fbf44550b7866201ec

Observation 32153b46-476a-4467-aeae-7b83e851470e · outbound

This paper cites Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.609868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.282744Z digest=sha256:59a51c6aea1f208400c9520d21c0e821e8ee0f540baf57d22a5fc8b149f7087f

Observation fc41d0e5-9d07-4f21-8904-5012e55c1384 · outbound

This paper cites Empirical production function free of management bias.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Empirical production function free of management bias

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.595983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.286681Z digest=sha256:83180bfaf581eb980ded95af2c511e4c2d7da69c31583c16879e67fd10cd73b0

Observation 56e8f4d9-725a-4e58-8577-1a4e69e10266 · outbound

This paper cites A Spectral Approach to Off-Policy Evaluation for POMDPs.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Spectral Approach to Off-Policy Evaluation for POMDPs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.290906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.290906Z digest=sha256:7e5fb4b8bbf091d901d9e7827575e063628c803c36ff5aef66a38424e190eb82

Observation 37f99efe-7629-4a6a-9394-86ee51e38913 · outbound

This paper cites Off-policy policy evaluation for sequential decisions under unobserved confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy policy evaluation for sequential decisions under unobserved confounding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.582136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.295005Z digest=sha256:4737e5af2a4cbd2399373ee133f3d93524eb670250b4b26e144a8f425cf3b972

Observation 54907508-141e-4877-a4e4-f4268c6c760a · outbound

This paper cites A Novel Embedding Model for Knowledge Base Completion Based on Convolutional Neural Network.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Novel Embedding Model for Knowledge Base Completion Based on Convolutional Neural Network

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.299285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.299285Z digest=sha256:fcc8b3b17b5221c1001c1e939fb595a36634c46120fe52ce47ed1b4f04f54c8a

Observation d57eeb98-a197-4cfd-92f9-518be18c6ab6 · outbound

This paper cites A review of relational machine learning for knowledge graphs.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A review of relational machine learning for knowledge graphs

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.567741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.304112Z digest=sha256:cb81c262dfad09d14e1170a880aff071c8491921bff3a62eb25be98888a7883b

Observation eca033d3-c3ed-4dac-b4a0-4093929bb8f9 · outbound

This paper cites Automated cars meet human drivers: responsible human-robot coordination and the ethics of mixed traffic.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Automated cars meet human drivers: responsible human-robot coordination and the ethics of mixed traffic

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.553282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.308697Z digest=sha256:ab074e330fadf2432fe456238b004fe0fd42022baa2282f37825654fc2ee65da

Observation 0502d09e-fd3a-4c68-99bd-16bd2ce1a339 · outbound

This paper cites blessings of multiple causes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning blessings of multiple causes

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.313162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.313162Z digest=sha256:4ae31a9da6971b16dac7e77433183dc11da68903c0197280ea06ac717b157e26

Observation cd511ddc-54a8-4ae9-8bf4-841a0c66b884 · outbound

This paper cites Counterexamples to "The Blessings of Multiple Causes" by Wang and Blei.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Counterexamples to "The Blessings of Multiple Causes" by Wang and Blei

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.317636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.317636Z digest=sha256:c151031538ae78f052dd2cc8a9979d0ff22889c080b45ae4fe9aabfa88ac9bb8

Observation f03a237d-00ca-423b-81c7-15fef16315ff · outbound

This paper cites A critical look at the consistency of causal estimation with deep latent variable models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A critical look at the consistency of causal estimation with deep latent variable models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.539284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.322630Z digest=sha256:933d98323607324d491cb4076ba90135c534bd66291eac5387703812d6f8d535

Observation 0df06961-fb28-404c-9c24-a728d438e613 · outbound

This paper cites Doubly robust difference-in-differences estimators.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly robust difference-in-differences estimators

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.525737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.327210Z digest=sha256:ae2436f32518c4787da6bf6c5eca1b4e4beb10640c8e34c203768897831b2e4e

Observation a4e00b95-682b-4952-895e-f771f7975623 · outbound

This paper cites Importance resampling for off-policy prediction.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Importance resampling for off-policy prediction

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.511247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.331634Z digest=sha256:4fadcc6e004b76f68d2bbbb0043ea12827d9f55f9498021b25592dad3dd5a077

Observation 76918f35-ec26-4eb2-a005-7527cb92a1ce · outbound

This paper cites Nonparametric regression using deep neural networks with relu activation function.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Nonparametric regression using deep neural networks with relu activation function

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.336054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.336054Z digest=sha256:60dfdfefa2384dad63a784dc3081ca0c301a712025dd95e34e85d080bf1be311

Observation c94b25b5-0eb0-49ce-b030-b6d9692cd316 · outbound

This paper cites On counterfactual inference with unobserved confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On counterfactual inference with unobserved confounding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:49.773072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.340724Z digest=sha256:02f24155d8ada54a3d45c88747bfa88c02c04fb02b25af53e6a4b9092e6248fd

Observation 067ac510-944f-4d18-b6cb-84dc2de34cb7 · outbound

This paper cites Does the markov decision process fit the data: Testing for the markov property in sequential decision making.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Does the markov decision process fit the data: Testing for the markov property in sequential decision making

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.495984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.345607Z digest=sha256:09e5e4d630859e6432f27f792bedb04955dd5e817b6641205c02c8c59fbda6d6

Observation dba602c9-a56e-4e02-af6c-fd7fbd4fcb16 · outbound

This paper cites Deeply-debiased off-policy interval estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Deeply-debiased off-policy interval estimation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.481440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.350052Z digest=sha256:7a5a1a718ffcb642304d8805b41fc323e8e22d3b73be8a5651cb405326af20f5

Observation 01a0abe7-28ec-40c9-b6fa-76afeb142d23 · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.467036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.354184Z digest=sha256:6a6935165d29ad606cce2d71145ec017d454f4fd2e1a54e8bff7a002c295167d

Observation 9397cc0a-09a2-462f-8e3e-0ea8651c291e · outbound

This paper cites Statistical inference of the value function for reinforcement learning in infinite-horizon settings.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Statistical inference of the value function for reinforcement learning in infinite-horizon settings

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.451972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.358732Z digest=sha256:0008bb0650fcf9804bc44c03bed86f8665358fe1218b6e28b99cc6862ad5a092

Observation f75aea1e-95db-497e-83dd-6c2ecff17798 · outbound

This paper cites Off-policy confidence interval estimation with confounded markov decision process.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy confidence interval estimation with confounded markov decision process

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.437357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.363172Z digest=sha256:a3dc80cf8c89242cc1a4da077a35efc444cda9bbbd34e75c78e23d40867aab0a

Observation 611f730e-97af-4690-94ad-937ecc2f2c3f · outbound

This paper cites Mediation pathway selection with unmeasured mediator-outcome confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Mediation pathway selection with unmeasured mediator-outcome confounding

Reference 79

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:26:49.752062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.367414Z digest=sha256:814be43b8fab99910e6b4bad8f453b37d356278e82cad232a1b23a6f6eed12c3

Observation 5a207d18-b0e6-48f3-9c1d-8d6dfc330435 · outbound

This paper cites Reasoning with neural tensor networks for knowledge base completion.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Reasoning with neural tensor networks for knowledge base completion

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.371847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.371847Z digest=sha256:ffe49e7dac999cff0a9adc45c4694a56c22ac609af439317060f11fc2993f88f

Observation 8bf13f72-5146-4824-b96b-f6c32cf73b4b · outbound

This paper cites Reinforcement learning: An introduction.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Reinforcement learning: An introduction

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.376356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.376356Z digest=sha256:5148e57184df58b7ef30c87e17927c61f1302fd28f9eaee195b876201a9fc09e

Observation e254b7c3-ddd9-4d3c-a61e-477b512f9f93 · outbound

This paper cites Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.381181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.381181Z digest=sha256:456125cf30f933dba161630c8007cc866c8f084cf1fd876f5baccaa72f130ca3

Observation ff5acfdd-eea0-4bea-915a-509ac1b12643 · outbound

This paper cites An Introduction to Proximal Causal Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning An Introduction to Proximal Causal Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.385666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.385666Z digest=sha256:ecd65ecaceba9c9c7c26837d858b89b17e6f3c8fdc973cbbd03ab72bf588fbba

Observation f53839b7-c0a4-4ed0-a2b4-dde6c539d20f · outbound

This paper cites Off-policy evaluation in partially observable environments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy evaluation in partially observable environments

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.405119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.390375Z digest=sha256:1eb518ad63ee9c362ba481646c11335052b757596aa3dd7710c85f4078e73bca

Observation 2b35f4a5-8d54-4921-9d90-dacbeb46223a · outbound

This paper cites Data-efficient off-policy policy evaluation for reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Data-efficient off-policy policy evaluation for reinforcement learning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.391180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.394915Z digest=sha256:6a15e078d1bcb13e90c00f5389b7b21a6ee9b1b15859b61875098ecad00197a0

Observation 3538c71c-a3a8-4ff6-975c-c9a846b1b822 · outbound

This paper cites High-confidence off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning High-confidence off-policy evaluation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.376590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.399443Z digest=sha256:e11a1e2951362b3bc5a525cbac82e67c170757c47ee11bfd30454131b8e3ee65

Observation aa3f4e79-c057-494b-94b1-6dd011a90948 · outbound

This paper cites Implicit Causal Models for Genome-wide Association Studies.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Implicit Causal Models for Genome-wide Association Studies

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.403988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.403988Z digest=sha256:135e2753537c9524ff726631a421803d4333617c532b80c5d6e003ca1293e6d6

Observation 8e91b53c-552f-45e5-ad6e-80d394d79687 · outbound

This paper cites Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.408990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.408990Z digest=sha256:c163b663479a6a6aeb8a3100e4dc971e737abf1f12db9fe642a6fc02baeb85ec

Observation b6b9adc8-d1cd-4b59-b49e-c42fafb9b975 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Minimax weight and q-function learning for off-policy evaluation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.413790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.413790Z digest=sha256:b5431ad97fb281acf5b7120cf5c9c5e58d4cb5a4b32ad0da5cbeeaecc290b329

Observation f501c8cd-c13f-4e91-ab43-6980c84cd4be · outbound

This paper cites Using embeddings to correct for unobserved confounding in networks.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Using embeddings to correct for unobserved confounding in networks

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.352098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.418982Z digest=sha256:e73a6ba1d1480a5a8e33d37118099d84329276c50411be24ff39fa208b8bbede

Observation da6374d8-9829-4a5f-b237-bcc3689aef74 · outbound

This paper cites Adapting text embeddings for causal inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Adapting text embeddings for causal inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.423736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.423736Z digest=sha256:c5fbe6cee2dfc38031a47e59f1538bef7af46de1a8fc5150d985a498858f5996

Observation 3b16c69a-bd9d-4fc2-a840-e839d600de9d · outbound

This paper cites Relational deep learning: A deep latent variable model for link prediction.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Relational deep learning: A deep latent variable model for link prediction

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.329080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.428544Z digest=sha256:1aacb642005f2e09185d252e241f12857ba06abdbf320d53fc79e440cd047eaa

Observation 1f3480fa-a383-465c-8444-211ed4b6ea20 · outbound

This paper cites Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.433207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.433207Z digest=sha256:7ac64196e2a9f5a47057c055388923f21fa09c3a2d079e47cbc1ffe1edf428cd

Observation 4b0b5083-cc16-4f7f-bcf0-a517aa9b41f3 · outbound

This paper cites Provably efficient causal reinforcement learning with confounded observational data.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Provably efficient causal reinforcement learning with confounded observational data

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.315863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.438145Z digest=sha256:2ff68a0d514a0a23b17b654b59fe00201c25071a22f25b0f8106ea39a39ab45e

Observation bdf0cca4-42c4-44f5-a381-04d74a0cc6f8 · outbound

This paper cites The blessings of multiple causes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning The blessings of multiple causes

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.302272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.442778Z digest=sha256:49da3b043a6eb61084b808b7be51684774afd3a19df678be6bf8971989702de0

Observation 62992a5b-55ea-4fb3-a59e-7beabeab7809 · outbound

This paper cites The Deconfounded Recommender: A Causal Inference Approach to Recommendation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning The Deconfounded Recommender: A Causal Inference Approach to Recommendation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.447182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.447182Z digest=sha256:d8c192f5036b1653633cabc365722005f95bf9576a14fdbe3ba1cb2480eb3ca8

Observation 358dc732-754a-4111-847a-086c856d366f · outbound

This paper cites Semiparametrically efficient off-policy evaluation in linear markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Semiparametrically efficient off-policy evaluation in linear markov decision processes

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.288702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.451301Z digest=sha256:0bf9d395de67c00795354caa67b236d4e8e1044a68437ff7d699a7a4a1840140

Observation 3c3c5a1d-a062-4258-96ee-c74991263fb6 · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.274909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.455826Z digest=sha256:55335cc6dc3af8feec5c2207013619ac8b14f5eeb1a072d97b5584bf4e6df583

Observation eb66cdb4-473c-43a7-a7f0-558b3da967f7 · outbound

This paper cites An instrumental variable approach to confounded off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning An instrumental variable approach to confounded off-policy evaluation

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.260239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.460191Z digest=sha256:31720447b02fe856c292a2caad85e236c00fd5ec88fc5249c4af1f2fabe21d91

Observation 7f363e52-78a9-4d07-b6c1-8d1bd082a64d · outbound

This paper cites Strategic Decision-Making in the Presence of Information Asymmetry: Provably Efficient RL with Algorithmic Instruments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Strategic Decision-Making in the Presence of Information Asymmetry: Provably Efficient RL with Algorithmic Instruments

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.464366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.464366Z digest=sha256:c9a078fc29564a48b14def6e309595bb81ea1d06253e71bd7e3a5941d5b8e137

Pith citing papers

Observation 76ef2f32-1274-4e25-b5f0-bf677d993731 · inbound

Training Large Language Models for Self-Explanation Faithfulness cites this paper.

Training Large Language Models for Self-Explanation Faithfulness Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-08-01T08:38:36.968420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-01T08:36:27.422590Z digest=sha256:2e70cb523aceec946d53a9815ef1ea7086756c9380b0c3e9728d3b0a92d8c90c