Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding

As of 17 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2505.00304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00304 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:54:14.609941Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd4073d0-0038-4952-adf9-63822c6ee9da · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.256109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.256109Z digest=sha256:4d0352d6d061ff110e4033e3ea23856adbddb04e7dda614d4dcbe0e33d2d7d1a

Observation 919d6f0d-5353-48de-8699-761051d85bab · outbound

This paper cites write newline.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.261809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.261809Z digest=sha256:ff8d435a4e0d6fbae70b37a828e921b747eee2c0acf0ca2fc41b0fb48c964388

Observation 410f8da7-edd5-4ad3-9709-f6359a423471 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.823487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.266879Z digest=sha256:7649b2e477503a518c9b51faa8bc64d6556fffd8c7d2bbc1fcdf9fd3e92e185d

Observation 6fb5d4e1-7c06-417a-ad6f-51111b68b2b2 · outbound

This paper cites (2008), Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path, Machine Learning, 71, 89--129.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2008), Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path, Machine Learning, 71, 89--129

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.806643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.271354Z digest=sha256:d90940f2d9b55ecd1996c12f5bd8fc017426862719833e6ded1e0ed11460b674

Observation 4e3b1f35-75ae-4d39-bac0-44115f865f80 · outbound

This paper cites (2017), Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research, 18, 629--681.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2017), Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research, 18, 629--681

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.790796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.275782Z digest=sha256:cff76f4b20e1ff09253ba2e4b6de2466d172406071b63c44034e32b3e593a28f

Observation 7b1ecb43-2fc0-4ad4-afc2-de323f09bd6d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.775627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.280602Z digest=sha256:b088271870492376031c5814c70c4ecc6d6a30d133a207d57a0289f146be26bb

Observation f3845a6a-eff6-4af7-b773-d37f0917cade · outbound

This paper cites and Kallus, N.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Kallus, N

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.760428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.285139Z digest=sha256:d88d0dd826a976593a386c147df6f1c41b3c3abe1f8a361c6939a338302ae597

Observation 2bfdfeb8-ce6d-418d-9481-3545b554ba51 · outbound

This paper cites (2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.745039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.289645Z digest=sha256:87de5d9492e5f33efea94e8c441056e761253b39f0075fede6867129dbc28bed

Observation b13a54ca-1684-4dbe-9b79-434d67a8905e · outbound

This paper cites and Kennedy, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Kennedy, E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.294412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.294412Z digest=sha256:4a4ae4f8ac5fb787f6db0f21094e98bdeb246d6dd9179f594a62f9bf1b3c2e3a

Observation 466524b1-2ecf-4ea2-b1dd-f7190268ce50 · outbound

This paper cites u derl, J., Schmiedeberg, C., Castiglioni, L., Arr \'a nz Becker, O., Buhr, P., Fu , D., Ludwig, V., Schr \.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding u derl, J., Schmiedeberg, C., Castiglioni, L., Arr \'a nz Becker, O., Buhr, P., Fu , D., Ludwig, V., Schr \

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.717702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.298932Z digest=sha256:088c8b8c788879187dcb95f4e19db2b49f231cddcb0899644f9ac1e665a11b10

Observation 11852540-b5e5-49c5-b25f-bf5ea33bb747 · outbound

This paper cites W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.702282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.303552Z digest=sha256:4edf35361368910a0ca94c227c8b78c6fd1177f08aec0cb31323f04c67b1b010

Observation ba310e6a-5471-49cb-849a-b1277b746cba · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.686011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.308223Z digest=sha256:4ad7f172392f6b260817d92d73198f4fd14112f61ab6d760048a6d179fb4a119

Observation 5ad80ee7-1129-4815-a00b-59980fb968fc · outbound

This paper cites Jump Interval-Learning for Individualized Decision Making.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Jump Interval-Learning for Individualized Decision Making

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:54:14.824601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.312790Z digest=sha256:5a02f040068dee323de111dfc18b67ac9a3a19cca811b8f12893b0a46d4eac78

Observation 818cf5dd-c8e1-4f46-a3c2-c31add3c0577 · outbound

This paper cites (2022), Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency, in International Conference on Machine Learning, PMLR, pp

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.670739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.317903Z digest=sha256:6aab994146c10a928371f1b4465180093cd2cc8ccc129e09bd13ed69d5344d34

Observation d2bf2f97-5c20-4453-ad01-9ca77c8cee3c · outbound

This paper cites and Qi, Z.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Qi, Z

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.654763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.322320Z digest=sha256:a471603cda1f2a092b5be1de9b0588c70c3fff19442db47d0900c28226e69ce2

Observation b0d9fe11-a3c4-40b0-8ad2-6ba8367df1f9 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.638815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.326721Z digest=sha256:9579bcb9b00b77cb662d9545a05ba3750ad0ba314b98bf8d26d9805dc637aa68

Observation 4e3b8511-e91b-4714-bcc7-1e34e89d429f · outbound

This paper cites (2023), Semiparametric proximal causal inference, Journal of the American Statistical Association, 1--12.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Semiparametric proximal causal inference, Journal of the American Statistical Association, 1--12

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.623843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.331138Z digest=sha256:0a1311d0df541c3e675990a60af27ada4dbf982713f511218ee3395fa14cac5f

Observation e9046a6a-e76e-46f2-93c4-d7e0f10f1179 · outbound

This paper cites and Tchetgen Tchetgen, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Tchetgen Tchetgen, E

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.335429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.335429Z digest=sha256:2f8d3eb4882ac383a8496aaaa7c0a311670859a9b18c1c08207da1c103eceaeb

Observation 29ef32ba-ce66-481e-a0ba-3e3a1705fd6d · outbound

This paper cites (2020), Minimax estimation of conditional moment models, Advances in Neural Information Processing Systems, 33, 12248--12262.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Minimax estimation of conditional moment models, Advances in Neural Information Processing Systems, 33, 12248--12262

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.597816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.339985Z digest=sha256:5fbb6c881982501489425f858359730732d112689dcae12a8e7f8424fa57f3a0

Observation 9867b87c-580b-44ac-83f4-0b8b48f2b823 · outbound

This paper cites (2011), On the completeness condition in nonparametric instrumental problems, Econometric Theory, 27, 460--471.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2011), On the completeness condition in nonparametric instrumental problems, Econometric Theory, 27, 460--471

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.583087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.344496Z digest=sha256:cf917e0aced806631dcf47010abeed90a1149202e06d0c1d2ec9315e73993e10

Observation ec68a9cb-3a54-4b1e-87c4-b4a7ba7c94a1 · outbound

This paper cites (2016), Regularized policy iteration with nonparametric function spaces, Journal of Machine Learning Research, 17, 1--66.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2016), Regularized policy iteration with nonparametric function spaces, Journal of Machine Learning Research, 17, 1--66

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.567752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.348905Z digest=sha256:c0d30a15e4e9a8eb2e42d2fd9ee6d17655ae0c70b11e4f77543762fa3b92d691

Observation d2ed938a-2311-4043-9762-36a090b59f61 · outbound

This paper cites H., Moreau, Y., Murphy, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding H., Moreau, Y., Murphy, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.551560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.353247Z digest=sha256:651aa847747a5decf53d3b200be1da726ce5a44e526961f0dc846921541157a8

Observation a4e09640-c797-485e-8822-c14d95c61cff · outbound

This paper cites Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.357740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.357740Z digest=sha256:519ecf797e9517566b66e265534940d517ce37eda0f37b1468840e4a5815a82f

Observation 97ab3828-eec7-472b-bb43-0f08ce7f9dc4 · outbound

This paper cites and Gu, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Gu, S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.536508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.362574Z digest=sha256:5877f9afbf944b743bb2fdc1be165ff710cd2298e399fe2181173c9a6d9e139e

Observation 22989646-d206-43d8-bf02-e108ba7e524b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.521397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.367001Z digest=sha256:f26c83732adf827dc15a658f2a86ac9c53c1f0e5a007278d673f2f8a8f028007

Observation ef9c2fa8-2725-4236-b75d-1b1b5b19a955 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.506348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.371400Z digest=sha256:021a5b0d415869ab09de54ffac2874958f15b71cb9d6380bac3e7563b8bff46f

Observation 7a388575-444c-4319-a6c5-0497d81ae671 · outbound

This paper cites (2022), Provably efficient offline reinforcement learning for partially observable markov decision processes, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Provably efficient offline reinforcement learning for partially observable markov decision processes, in International Conference on Machine Learning, PMLR, pp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.491679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.375883Z digest=sha256:aab644c7f09031b91658a4c660b88464a99bef304e809c21a6b5d2b976e1071f

Observation e918f253-246e-4e1c-9c86-be1d7007cfde · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Soft Actor-Critic Algorithms and Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.380110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.380110Z digest=sha256:f5d8aac7949d75d4bc447168203edfb38ccab984d30329ccc9ab6f88aed39172

Observation a0c36dbc-7f04-4b38-b8ec-feabeff990fe · outbound

This paper cites (2021), Bootstrapping fitted q-evaluation for off-policy inference, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Bootstrapping fitted q-evaluation for off-policy inference, in International Conference on Machine Learning, PMLR, pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.476082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.384888Z digest=sha256:9412f97c40a4cd924610acf3d942c9d53dd435e87fd9eeddb63c3a9568438338

Observation f542dcb4-c1a5-4151-b49c-b8e10f1fc051 · outbound

This paper cites W., Lazaric, A., Ghavamzadeh, M., and Munos, R.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding W., Lazaric, A., Ghavamzadeh, M., and Munos, R

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.461072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.390238Z digest=sha256:526ae80068eab862adb3a46991edbb010522c096c685a4403dcdb6ad1698b894

Observation 2c93aa95-6b44-413d-b947-93e017b98bb9 · outbound

This paper cites A Policy Gradient Method for Confounded POMDPs.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding A Policy Gradient Method for Confounded POMDPs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.394823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.394823Z digest=sha256:464bae133afeb11982c63474063d954f311311dca55b6e31faa453056b178f16

Observation 4c83a170-793b-4e33-9c36-92dee9e03d32 · outbound

This paper cites (2020), Sample-efficient reinforcement learning of undercomplete pomdps, Advances in Neural Information Processing Systems, 33, 18530--18539.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Sample-efficient reinforcement learning of undercomplete pomdps, Advances in Neural Information Processing Systems, 33, 18530--18539

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.445717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.399509Z digest=sha256:6f2a2b171a24d8470ae9903437f6ed916948f6cf9f5cd7742b19b13ac5461269

Observation 679dbfa6-731c-40af-aaa2-0efcd270d2e8 · outbound

This paper cites (2022), Doubly robust distributionally robust off-policy evaluation and learning, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Doubly robust distributionally robust off-policy evaluation and learning, in International Conference on Machine Learning, PMLR, pp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.431262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.403922Z digest=sha256:476c6dcdedb19abe553801af6ae2ae93aebe76a085aef54f4d89d0452a28b720

Observation 41912fe6-7a55-40dc-bba5-46b1dde2d57c · outbound

This paper cites and Uehara, M.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Uehara, M

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.416571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.408243Z digest=sha256:a4758c1d16798dd84f9e5b27f1177363bc1b19572e319b47ed2a03400a2f7f1a

Observation ad1d724f-0a8b-40d4-a0f1-338d96381689 · outbound

This paper cites and Zhou, A.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Zhou, A

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.401782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.412788Z digest=sha256:137770a6094eec38de1604c53c6ccd7c804c4a72db780f8f97a3e1026a3217df

Observation 770b2abc-3d1e-4920-bab1-b7e6034d2b1b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.386488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.417271Z digest=sha256:92be15a640cf4fbc10b9452c3f405d868161fe0a1cc47408684865626767d01e

Observation 86f2fbbe-2e73-422d-82bb-c9ee1b033d39 · outbound

This paper cites B., Volkmann, J.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding B., Volkmann, J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.371758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.421631Z digest=sha256:8c2bf3732ff91df45383a22000ec6ec7905c1da42a62ea47aac44b2f934b80e7

Observation 841944cf-839a-4b90-9e6d-a176aa198162 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning with Implicit Q-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.426021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.426021Z digest=sha256:2b76c6dfc8a5f4609488f2e805611c3fd6ba0469caa4ce8a513cb1fc9bb72ff8

Observation 8864b3c9-f08c-4ef0-a21a-8980cb3ed206 · outbound

This paper cites (1989), Linear integral equations, vol.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (1989), Linear integral equations, vol

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.357186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.430679Z digest=sha256:206b645a3deb5d5ebea880b7569548788e54c40b18d5648debf95a491a935dfb

Observation 03655345-0d94-4581-9c58-68eb06c787cf · outbound

This paper cites (2020), Conservative q-learning for offline reinforcement learning, Advances in Neural Information Processing Systems, 33, 1179--1191.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Conservative q-learning for offline reinforcement learning, Advances in Neural Information Processing Systems, 33, 1179--1191

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.342241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.435075Z digest=sha256:1cf5e836235ab624cb3c2eda31e56a8ffd430baad03e2d0d25db61e80fb7c73b

Observation 8849dbc8-95c4-47ae-9032-6fe596954bd0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.327486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.439297Z digest=sha256:bce8da3f768e0c21efa4f7bc2846246d0d21099579069e746c07d57410fba18d

Observation f88e3568-60d5-4d53-8989-91dbb3c9016c · outbound

This paper cites (2019), Batch policy learning under constraints, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2019), Batch policy learning under constraints, in International Conference on Machine Learning, PMLR, pp

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.312552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.443453Z digest=sha256:cc8f480eab9d8e20c8d310df9272709c03df1236da1307a2cab3769f02da568d

Observation b7f6d36c-6ed5-467d-a10f-e67b1502c283 · outbound

This paper cites (2018), Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling, in International conference on machine learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2018), Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling, in International conference on machine learning, PMLR, pp

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.298204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.448080Z digest=sha256:7769914e5dc8cd20da0734e3e33c8fe35cfc8ed21473f0e66939ed71dbf16973

Observation 8048a3b7-f55b-4017-9e54-1f6617564eb2 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.452593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.452593Z digest=sha256:347605988279c736efc184010b1e3e340da8db0f3ef1793ef53979919a68ea27

Observation 14bd91e6-4c14-4b02-9245-b2dc72627b3f · outbound

This paper cites Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.457368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.457368Z digest=sha256:76457f4ad6c98072fd096c2786fff22fd41b395fe2e6695a303ecc75c1af0606

Observation 70bb7568-6d3e-4162-86a0-2cb47449696a · outbound

This paper cites (2023), Quasi-optimal Reinforcement Learning with Continuous Actions, in The Eleventh International Conference on Learning Representations.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Quasi-optimal Reinforcement Learning with Continuous Actions, in The Eleventh International Conference on Learning Representations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.282669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.462291Z digest=sha256:304ee56b9aecbf6f061729caaf6a6f603f65c68948ebb043d24d9d027169b977

Observation d0436792-f256-4aa7-bdde-4aeb45b17928 · outbound

This paper cites (2021), Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association, 116, 382--391.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association, 116, 382--391

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.265633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.466677Z digest=sha256:0827dd860916b27c7fafef35c403aa0cfa3c32908067f82a10229de4a20d5b33

Observation af9514e0-1a5c-4152-88ea-09e1ae96af7c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.250381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.470965Z digest=sha256:17f6ce58de27fe20cc92b8bf9ad7dbfc8850d891316da53b5dfe09d984961182

Observation 73b4f8ed-f353-4222-9a7e-006ea9ebf242 · outbound

This paper cites Continuous control with deep reinforcement learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Continuous control with deep reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.475241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.475241Z digest=sha256:4e6bf728ba0e9759b0f80994a1856d102735bbcf57220d4939143e908cbee6a5

Observation 0e94eefd-784f-4d04-9bfa-9acd74d2424f · outbound

This paper cites (2018), Breaking the curse of horizon: Infinite-horizon off-policy estimation, Advances in neural information processing systems, 31.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2018), Breaking the curse of horizon: Infinite-horizon off-policy estimation, Advances in neural information processing systems, 31

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.234853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.479814Z digest=sha256:858f5042a3971102481431f915c7bedfbfc7b77fab5afb3aadb510a535b29758

Observation 2a5f2215-1b5a-4ce8-9779-ade6115e83c1 · outbound

This paper cites Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.484531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.484531Z digest=sha256:fa212de885b10b23f7f3405f5210aa18fa9757d6c90a6393ca94726c93fdc3b7

Observation 0f8924a3-0a2d-4df5-92b6-271946ed8cb1 · outbound

This paper cites J., Laber, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding J., Laber, E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.219656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.489344Z digest=sha256:17dd3536279091129c64151c8575e8eb63564a94b9cef49247df971ecdf31bc6

Observation 2d7fbf26-4a78-4e5b-9105-730ec6aad78d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.203969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.493813Z digest=sha256:148e535c500453eba3b35e80695913a4be7712a1e1325136abd9469f48baca4d

Observation 356b302c-a8b7-4be8-aa11-789b56f312a8 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.498384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.498384Z digest=sha256:0c7ec54a097d43caa11076812e829a6dc2c9876bf3ac071461058233805ca817

Observation 93f237ba-a115-49e9-98e9-9a4756e8fc95 · outbound

This paper cites A., Veness, J., Bellemare, M.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding A., Veness, J., Bellemare, M

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.179126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.503080Z digest=sha256:17d3cb3399d93dc87e2ce56bd9f3b826c4af596e2c7c6d9eb6947c7c67363cd1

Observation c198ad41-3aaa-4cbb-9fb4-b77187b50c60 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.164405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.507453Z digest=sha256:fd45b24d9b9616db7da8a34d61420120f198466b96bf594d9f0d763b0f03534b

Observation ea625b0f-788b-4480-ac5c-fa78767a645b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.149351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.511795Z digest=sha256:49269b4d623fcd3679fdf59ec696478cd65c47bb26a9ca1fad54b2d6ffe7ce2e

Observation e4a26aa4-66a6-4696-8ef2-d0abb2353646 · outbound

This paper cites (2000), Eligibility traces for off-policy policy evaluation, Computer Science Department Faculty Publication Series, 80.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2000), Eligibility traces for off-policy policy evaluation, Computer Science Department Faculty Publication Series, 80

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.134256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.516271Z digest=sha256:79494e56a9c751b999bdb527773518838a8639ef7eef4ea3814d4f36f77014cd

Observation eb3bbccb-0449-4776-98fa-674978f528d4 · outbound

This paper cites (2023), Proximal learning for individualized treatment regimes under unmeasured confounding, Journal of the American Statistical Association, 1--14.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Proximal learning for individualized treatment regimes under unmeasured confounding, Journal of the American Statistical Association, 1--14

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.118195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.520939Z digest=sha256:cf2982c6cf049bb06720bc41741bb9421f8cbc5821c578a2a0411506a93f6674

Observation d3d81530-ba02-40af-b8b8-e5fa33af6952 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.102803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.525416Z digest=sha256:5b02aa4cb778c510feb5dcaf9bfc00685291d1fcf8b1b2fa6ab5640d08326593

Observation 97084137-1bcb-48e2-900e-7e0333532a06 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.087809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.529849Z digest=sha256:018fce93d5f6e3f7e39ef22f539541ad551d22119f0a04602dc2897535803be2

Observation 07b18144-1ed5-4b4a-9859-895647d7f1fa · outbound

This paper cites (2022 c ), Off-policy confidence interval estimation with confounded markov decision process, Journal of the American Statistical Association, 1--12.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022 c ), Off-policy confidence interval estimation with confounded markov decision process, Journal of the American Statistical Association, 1--12

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.071842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.534202Z digest=sha256:34678b360d8b058e51699e2a5ba1d9587ad5c59cc1b420c36338ca46e3a61b9a

Observation 8eb2c479-d3e8-497c-8163-4f098ac330b4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.056559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.538514Z digest=sha256:a134e2a0c3b0e8a444b48dbbc412c58d41f464bea43e065293ac1deda71a0c50

Observation a206f1f6-67d3-45de-b359-cffd17e5a38e · outbound

This paper cites An Introduction to Proximal Causal Learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding An Introduction to Proximal Causal Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.542761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.542761Z digest=sha256:025abb5684b332e15ea4add391451c6a8d1a7b9c5868eab48e0701695c993f58

Observation 0e6742b5-5241-4b4f-b984-ea5e1d6c44f3 · outbound

This paper cites and Brunskill, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Brunskill, E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.037610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.547332Z digest=sha256:251c501a43c9acd5414ede5ed97d7865dcf38883ebe916e5c4d699a8d21ea5d4

Observation b22ce68a-5aad-4f34-a0d5-285f3cf0ffdd · outbound

This paper cites (2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.020578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.551671Z digest=sha256:27d5041e32a81aded131f9b3ad2359ce546ca69fdc31a3955752647c72a06bff

Observation 0b3f6020-ec07-4882-a829-8ed66983dec8 · outbound

This paper cites (2024), Future-dependent value-based off-policy evaluation in pomdps, Advances in Neural Information Processing Systems, 36.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024), Future-dependent value-based off-policy evaluation in pomdps, Advances in Neural Information Processing Systems, 36

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.000899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.556461Z digest=sha256:f85d7fa1f03365ce6b2377d782e30e0d9e1833b06d39ab23c8ee7c73c739c056

Observation ed0aa84a-7f48-49c7-993a-2d307ef8202c · outbound

This paper cites and Groothuis-Oudshoorn, K.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Groothuis-Oudshoorn, K

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.985762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.560704Z digest=sha256:75aee0ecbfcdd26c86fa0f820e35f972db8b603b7989377024569c0ce66d5902

Observation aff82c72-59d1-47d5-907b-b6d4222f706c · outbound

This paper cites and Zou, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Zou, S

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.971136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.565105Z digest=sha256:a2390ad7e6f8e731e9d1b2e8e34085d73f8fbc113e99c847b5bba5f4d228ff31

Observation 3e1da517-73bf-45f0-a826-4367b3fdb266 · outbound

This paper cites (2019), Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling, Advances in neural information processing systems, 32.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2019), Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling, Advances in neural information processing systems, 32

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.955283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.569342Z digest=sha256:3ec233a7acfc6deb27e6fb1b787c21e076f404675b14d2fe8710e5f279f28fad

Observation 0a1d9ce9-2d9b-4746-8d99-b9c128bb581c · outbound

This paper cites (2023), An instrumental variable approach to confounded off-policy evaluation, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), An instrumental variable approach to confounded off-policy evaluation, in International Conference on Machine Learning, PMLR, pp

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.939449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.573726Z digest=sha256:3027d1a437c7ea1dd052efa7b9a2bb9ce060301053f4212dc10a232571679f57

Observation b62d42a6-c9b9-4314-a023-203d356b22a9 · outbound

This paper cites and Bareinboim, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Bareinboim, E

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.924013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.578035Z digest=sha256:e7d9d62533c24bc7c9c69e44d6d438f0c5605dc9dbb0fbacbdc32da2d190e4e7

Observation 686d4d10-0ba5-4f96-a399-6911a249dcc7 · outbound

This paper cites (2020), Causal imitation learning with unobserved confounders, Advances in neural information processing systems, 33, 12263--12274.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Causal imitation learning with unobserved confounders, Advances in neural information processing systems, 33, 12263--12274

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.907724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.582444Z digest=sha256:9c299b31e32aed8602c9f1de6baab38873c14ce1a368b0b68073bc56b4b84c7e

Observation 98108e71-ba89-4819-8277-83b9d4093515 · outbound

This paper cites On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:54:14.671602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.586780Z digest=sha256:e7da1c987ace1da8b79ba67bbf182eaa8bdf82fd8343ff24a9100ee95972802d

Observation a9b94b24-e0ac-4371-8991-def5ab75dc3b · outbound

This paper cites (2024), Bi-Level Offline Policy Optimization with Limited Exploration, Advances in Neural Information Processing Systems, 36.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024), Bi-Level Offline Policy Optimization with Limited Exploration, Advances in Neural Information Processing Systems, 36

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.891413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.592125Z digest=sha256:09a3a5520d41de8af593edf3385994ec223113a88fe5c6d277162df3509a0c8b

Observation 8d09bd4e-b1cd-40e3-9026-7d253601b194 · outbound

This paper cites (2024 a ), Policy learning for individualized treatment regimes on infinite time horizon, in Statistics in Precision Health: Theory, Methods and Applications, Springer, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024 a ), Policy learning for individualized treatment regimes on infinite time horizon, in Statistics in Precision Health: Theory, Methods and Applications, Springer, pp

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.874788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.596573Z digest=sha256:a7ada0899749cc6341643043f86259d0f4efc1ccd5602336df72538669963af1

Observation 5c19e899-4c9a-4945-8ca5-103002d52080 · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.601000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.601000Z digest=sha256:d15c9ef94c018697f9717473682b9fe322c3a123ec4656dd678483b40bd5673d

Observation 6fe62085-e851-42d9-827b-f2b8ace0bd17 · outbound

This paper cites (2024 b ), Estimating optimal infinite horizon dynamic treatment regimes via pt-learning, Journal of the American Statistical Association, 119, 625--638.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024 b ), Estimating optimal infinite horizon dynamic treatment regimes via pt-learning, Journal of the American Statistical Association, 119, 625--638

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.857615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.605462Z digest=sha256:6c59fbb98d2134076b811cdf823fb98062cd7921a62241053c9d397c6ce22356

Observation 7798237d-f28a-4154-8fe1-f5271244b8da · outbound

This paper cites (2020), Safe, efficient, and comfortable velocity control based on reinforcement learning for autonomous driving, Transportation Research Part C: Emerging Technologies, 117, 102662.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Safe, efficient, and comfortable velocity control based on reinforcement learning for autonomous driving, Transportation Research Part C: Emerging Technologies, 117, 102662

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.841403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.609941Z digest=sha256:45b206cbc6ae7ad304a462b575b03ce0667fd1abe34f2334585de54b370d35dd

Pith citing papers

No inbound Pith citation observations are available.