Pith. sign in

Paper Citation Record · LEDGER

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

As of 16 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2505.22492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22492 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:17:01.124208Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact7
  • verified fuzzy50
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe970c9f-0548-4569-aaa3-4eab2ff15df9 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:52.903422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:52.903422Z digest=sha256:cea6a63a578085f50e8dd97d1f1729759971a6705c172476c07b0817c105bf77

Observation b62aa258-e78b-4938-bb97-799ce524f0ef · outbound

This paper cites and Kallus, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Kallus, N

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.021013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.021013Z digest=sha256:bb9314eff6d0fe646646ee4d1d87238ab17449bf878f9f1b67b46bccac4d2320

Observation 3b741a11-b4dc-4569-b70d-945c5839b408 · outbound

This paper cites Off-policy evaluation in doubly inhomogeneous environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in doubly inhomogeneous environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.103572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.103572Z digest=sha256:202d12aab598f3d66585c2e493ad1bf5f75cada7abad687f4b0b1f3c9f26505a

Observation 440371c0-981b-4670-bec6-f0aacf55fdae · outbound

This paper cites More efficient off-policy evaluation through regularized targeted learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More efficient off-policy evaluation through regularized targeted learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.186263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.186263Z digest=sha256:fd8ab1c067ef408c7b7617d4f994a3ebac7d33f967d4789298a94b276ddd8bcc

Observation 342ab1ad-17f1-455a-9f63-4a19cc565efe · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.264406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.264406Z digest=sha256:312fb1c2d9df4707bb2ab834caf58e890acf96aa0358608d983185c62555ead7

Observation 552f657a-6eda-4337-a6d5-a218da5e2d98 · outbound

This paper cites OpenAI Gym.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation OpenAI Gym

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.388701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.388701Z digest=sha256:92e7c298b128b7129436bbf90aef0aeb08d043c9e3e6cda95a168d1d745d3ff9

Observation c6c5e743-b5b3-4cbb-aebf-ad81d8d95b29 · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:17:04.709883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.496289Z digest=sha256:0a68c00fdd1c03219837261ac3c432873fea9f02192d61ad67548a10935d0beb

Observation 8a06d0b8-b9d7-40c3-916f-86b10aa7c50a · outbound

This paper cites Structured Difference-of-Q via Orthogonal Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Structured Difference-of-Q via Orthogonal Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.447238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.610687Z digest=sha256:56cb26559d003ce27048d128dc8e4a9b28109f8c91b4ce4730ab5b25affdcc5f

Observation 67a52a5d-4608-4582-a6c2-bd0f88c6a259 · outbound

This paper cites and Berger, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Berger, R

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.686598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.686598Z digest=sha256:1fb88162295088c09266514e35c4e854f388e4969a2b4289157af4e5f4d84a48

Observation bc2ae633-0ad7-44ca-a0db-07a0658b0efa · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.794136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.794136Z digest=sha256:4e43c792005732eecccac6691ddfd64b84988f402efadc9710cfcdd7fa4b7510

Observation 49de3aaf-d1f8-4dbd-a4a4-759ad9b84b90 · outbound

This paper cites and Jiang, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Jiang, N

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.924369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.924369Z digest=sha256:11dfdc8bb352b509675035b0dbd5fc9f2f5d88f5873698c38b58b94371bdef08

Observation b8a8b783-825e-4208-95fc-5d209ea87241 · outbound

This paper cites and Qi, Z.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Qi, Z

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.023087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.023087Z digest=sha256:513e261b65a3d5bce6e94fe2fa5bf5e22ab6c0bde4e0a5e2071e06922424129f

Observation fca476a0-cfa0-4d92-ab2b-9bf5fb2ca2f7 · outbound

This paper cites Gaussian approximation of suprema of empirical processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Gaussian approximation of suprema of empirical processes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.133862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.133862Z digest=sha256:d0f319f4adb84a127be652102ab3262c3680d98f953bdccbb4fda61bca67cd58

Observation 52b4bb87-f570-4feb-84e1-1403a8ab136a · outbound

This paper cites Double/debiased machine learning for treatment and structural parameters.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Double/debiased machine learning for treatment and structural parameters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.246149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.246149Z digest=sha256:24736ee5afdd117ed7652c5b0c60e8c8019ecf3358b0a6a465691b7161cf6f0e

Observation db63a9a8-f309-47aa-acac-bd6bb2d1c379 · outbound

This paper cites Coindice: Off-policy confidence interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Coindice: Off-policy confidence interval estimation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.339217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.339217Z digest=sha256:2aafcd089971fc27175fbd1eefd9462c20f7ea1f71062521fd3f8b35ef7d1c9f

Observation 0a9c8274-3776-490b-a0ed-558411d6bf08 · outbound

This paper cites Doubly Robust Policy Evaluation and Optimization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Robust Policy Evaluation and Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.432077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.432077Z digest=sha256:07aa32c36ee1f7928b70ee3d728b5edef467cdd5c912f92127edf1891a759893

Observation 8446019f-3832-4df8-9d51-6b1368b142b9 · outbound

This paper cites A theoretical analysis of deep q-learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A theoretical analysis of deep q-learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.524675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.524675Z digest=sha256:cd940f19aed8365f8d5864537ffc5a9d02b7a43eefe7ede615415660d8352c35

Observation 25caa724-aca7-4a07-b9f0-dda5c47cd091 · outbound

This paper cites More Robust Doubly Robust Off-policy Evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More Robust Doubly Robust Off-policy Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.617714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.617714Z digest=sha256:137912d2fa2561032d786bc2a1382bedc3d27c032cf003e24e3c5e1c7c5522c6

Observation 69d160b1-511b-4984-a4d4-ec8006edc122 · outbound

This paper cites Accountable off-policy evaluation with kernel B ellman statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Accountable off-policy evaluation with kernel B ellman statistics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.732514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.732514Z digest=sha256:34f47342b5baaca38ca27a5fcfe2e3819c5a1380ffbe504c7fe520910b83a200

Observation 1d2f3fc2-4499-45b8-ae5c-0061bd1e4779 · outbound

This paper cites Combining parametric and nonparametric models for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Combining parametric and nonparametric models for off-policy evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.828722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.828722Z digest=sha256:88657d730d2ef9ae8ae417b56d4ff0b7a13c9ea2594ae49822076ced2be46c1d

Observation 3e674bf6-aede-4609-a97b-8498caf0fb30 · outbound

This paper cites D., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation D., Thomas, P

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:15.124519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.910774Z digest=sha256:b8493a421e4dc54b4356a691aa712dd680dca5ef06e5cc196220476cbacbd811

Observation 238d6f10-29ef-4d5b-b3bf-b06331b78393 · outbound

This paper cites Importance sampling policy evaluation with an estimated behavior policy.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling policy evaluation with an estimated behavior policy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.980638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.995797Z digest=sha256:602360ee4110a1557824171fdf73616765c30c36faa1d3d4ab4c3c7bca41380b

Observation 715be3ee-ce98-4c84-ac12-c5babdbdda54 · outbound

This paper cites P., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Thomas, P

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.829766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.071207Z digest=sha256:1c600af3ae556d87298babcca0ebe7c6c30df46f1f0ca5e90706ae4948ab560d

Observation 5db9c47a-84a6-4898-9dc1-e9c7a05cdcd1 · outbound

This paper cites P., Niekum, S., and Stone, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Niekum, S., and Stone, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.679690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.196586Z digest=sha256:e32860a09fadab01b3cf69103ac72378d0039df301dc44fc43c818bfcf559240

Observation 9c90fb45-6232-4e37-b518-1c03af849c8c · outbound

This paper cites Bootstrapping fitted q-evaluation for off-policy inference.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Bootstrapping fitted q-evaluation for off-policy inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.525737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.317248Z digest=sha256:7465f1bc0de90536b79ab3e20501f202d5de5981d98f0173b59472cfb167d1b6

Observation 891b15a5-8bf1-45dc-abba-88a38f30a1b1 · outbound

This paper cites Importance sampling via the estimated sampler.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling via the estimated sampler

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.384125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.387170Z digest=sha256:96fb559b7b55778e509f926cde66c16485378998638788f2e26f217e529baf75

Observation 2ad6d719-9b29-446e-a5d2-f9e1a16fc9ca · outbound

This paper cites W., and Ridder, G.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., and Ridder, G

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.222790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.459117Z digest=sha256:85bc1203f55829086ed39c3732c553dd8a7319a597f8138c8ec5b869d1a975e5

Observation db7f3592-4268-4450-ba42-a0663dfe4af2 · outbound

This paper cites and Wager, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wager, S

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:55.534857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:55.534857Z digest=sha256:b8024c07b2a97c41214a72ba942ca4d59536630d0feffe41737f605c2d51f177

Observation ff3d687f-0b80-4dd4-a4d7-db5b0d93866d · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.029859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.616369Z digest=sha256:d59d2b4d5f28a2621fb42ba8d23c06b7d05d4185e0cb54057b11ec5ab79728dd

Observation a42c0dd8-5648-4354-b424-84d1a4e77faa · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.838362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.716241Z digest=sha256:4c7afb30d934b5589970efaee270398e5a2c2eb65d0d765af67d331de0ab3db2

Observation d404d600-7de2-4150-946b-5307bac95f8c · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.626828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.793014Z digest=sha256:378e14df68d8fd173777281f4e20c974e40891ca6e75fb93bb685304ef663cae

Observation 6411d0f2-b200-46ad-b612-6dd15b17e52c · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.420849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.872168Z digest=sha256:ae04d0ad01cf20f0d64fce106809c1b1e9c1c6a3c7b90f773239edebef430d65

Observation 63221bff-3c62-45ab-9178-79969858e05b · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:13.160576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.932696Z digest=sha256:24784f671ed12c158d32ee2c7cac8d05c809e0dc8fa1c6eb76bd5b79b6e2b20e

Observation 5a9cc8b5-daf7-428d-9acd-b4d241a9daf0 · outbound

This paper cites Batch policy learning under constraints.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Batch policy learning under constraints

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.937701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.017006Z digest=sha256:d1990fed05f2534b3e1b69ac688546e696fa78346520d640778ac403175ffb3b

Observation 1924c1ed-bc7b-4f1b-98f6-816e702217d1 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.144780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.144780Z digest=sha256:079a2affa0a1b9f6fd58b1f36e2d95c265cbf0f541a6b44a7d7d77e6392a573e

Observation 977884f7-69c5-4e80-aa98-3b38ff7c87f1 · outbound

This paper cites High-probability sample complexities for policy evaluation with linear function approximation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation High-probability sample complexities for policy evaluation with linear function approximation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.262424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.248494Z digest=sha256:63074abbef8690658d7f609d1175fc65fd1bdfae614a9968f375befc258a6e6d

Observation 38ccf486-2bf4-47ce-91d8-e2cb5ef3fff3 · outbound

This paper cites Off-policy estimation of long-term average outcomes with applications to mobile health.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy estimation of long-term average outcomes with applications to mobile health

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.773972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.359148Z digest=sha256:55029d98df79ca748c47cd8b086a4e1859900c9c2440604925bd917f82ab1e68

Observation 67f38c91-6c96-4360-9f51-f15cbc3b435a · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:12.624851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.466433Z digest=sha256:5e607b646f899661e481ec5a857f1dc686e66f75fad87122e6f3df6fc2c1433e

Observation 03b93b2d-2ad1-411f-9a43-8b252edd6bf6 · outbound

This paper cites Breaking the curse of horizon: infinite-horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Breaking the curse of horizon: infinite-horizon off-policy estimation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.478362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.565112Z digest=sha256:60d6adbe4049a8a58e7364f4bd06b745bb20b7e21630090cb443da2100535c7b

Observation b1b84187-62f8-49de-bd51-6289711aa20c · outbound

This paper cites and Zhang, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.254767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.671738Z digest=sha256:14b2f1a097d77354a103d1f7cced596e5a5c6a23816209f46c408f3d212a9b31

Observation 20212785-c506-4558-a002-57494379b523 · outbound

This paper cites Doubly Optimal Policy Evaluation for Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Optimal Policy Evaluation for Reinforcement Learning

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:17:03.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.740396Z digest=sha256:9a51c83269669bc32eaca2742043141b085462cba0e714aaf9eb98f8f5816590

Observation b6029ff5-5b9f-48b3-9947-4d1d444ba384 · outbound

This paper cites Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:03.551774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.834090Z digest=sha256:e62a93667a51fda9d7c33cab7429233b9713cb8e07b9416777a2609913a632cb

Observation 1ed53d7c-3231-44c0-8f6b-0feeedbf69dc · outbound

This paper cites J., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation J., Laber, E

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.946901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.946901Z digest=sha256:00f7a92558b5b36e9cf6b2a27a1742a68fce2cee18ee34d7d59c7a9fac2e45e2

Observation 548001c1-8ac1-463d-adbe-e8cb79d621a4 · outbound

This paper cites P., and Nowak, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., and Nowak, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.077014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.007129Z digest=sha256:2c9f4cf9117aded379ea255ba338e9de976c07c23a0c252d1c158676abebc6a4

Observation e038af15-6e2e-40e6-8d48-d63427dd26d3 · outbound

This paper cites A., van der Laan, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., van der Laan, M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.923588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.074159Z digest=sha256:1196bfa2bc3a7540c39941408c0b11d50eb718ae9ce5ca1b94e9132b73e72306

Observation cb654943-6911-4027-b699-509c2785a743 · outbound

This paper cites Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.691191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.135333Z digest=sha256:1671d5d7775892b06910cdb9ee0bf9fcee50f44286ec1f5712069189942b2d55

Observation 3bbf4a0c-7e88-437c-ba2e-7be99f60802d · outbound

This paper cites A Spectral Approach to Off-Policy Evaluation for POMDPs.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Spectral Approach to Off-Policy Evaluation for POMDPs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.195687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.195687Z digest=sha256:8d9e963c2ca7f7c05fbf6e129c09e2373d7a45fde8b1fe8b1cc0a89b23b85f05

Observation 877ba37c-7916-41a1-bbda-3021e7f63f2b · outbound

This paper cites Off-policy policy evaluation for sequential decisions under unobserved confounding.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy policy evaluation for sequential decisions under unobserved confounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.540247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.246652Z digest=sha256:19232d8a2de00e8f83bf33a63057e90863870e24a882c9771c272d5fc33f230e

Observation fe049d9e-3dad-43a7-876b-39568be90b97 · outbound

This paper cites K., Hsieh, F., and Robins, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation K., Hsieh, F., and Robins, J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.289965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.316362Z digest=sha256:29004bb5aef71e0ce87883d05efb76e1f33bacda71e83d9bd8bec382006cc745

Observation 220033ba-5dc0-445d-b8e8-855b2dd8aecf · outbound

This paper cites Training language models to follow instructions with human feedback.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Training language models to follow instructions with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.362862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.362862Z digest=sha256:a0912851fbed462f2c1c63d0c1d4d55e9368bfeb17372f258fac05fba864c2ba

Observation 7e730401-bf8b-41ff-a17e-a81d69418f46 · outbound

This paper cites S., and Singh, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., and Singh, S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.104301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.449604Z digest=sha256:f738537d430f6a215365249961a70a97b3fcff579f28d3bec7fac6a2f4b94931

Observation c5644559-bd83-4111-bc92-4c2224a768b3 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.519872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.519872Z digest=sha256:cfc8f46a60f91e15f14894dff3049b87ec38f363a3cce3e10f58e9527c3d22f7

Observation eb28c104-9602-4e69-8561-9b0715d8a30c · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:10.808173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.582139Z digest=sha256:a949120080ff475fe5d9b974c05a465d63a19b3800b1140c947748418c71c652

Observation 4695e73b-5847-4e3f-82b9-e682b16dc01b · outbound

This paper cites Conditional importance sampling for off-policy learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Conditional importance sampling for off-policy learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.692678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.632221Z digest=sha256:cca610b92fc994410d1b09c6898379af7876f43c5640c0b09fca90af74113c28

Observation a5078da9-a7fa-4985-85d6-b91f6d93ef28 · outbound

This paper cites G., and Dabney, W.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation G., and Dabney, W

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.457993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.718028Z digest=sha256:0dfd204df11273a2f82aaafe7e75d3dcac37a79f05f282190fecb12699315389

Observation b9e6c4a0-196f-4998-8a2a-b23b71887564 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Proximal Policy Optimization Algorithms

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.801762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.801762Z digest=sha256:cd0eae7c600c7bc6fce44b23b6e54bd1f82a8d89818d168ef19195b609e9613f

Observation 7e7a7233-70a3-4e87-b622-49670b846ed1 · outbound

This paper cites Estimating the dimension of a model.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Estimating the dimension of a model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.220070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.866226Z digest=sha256:0f57b284e0639d051cc4b02fa2f15d3a8c70ef6fe45ca3a52af152e5e0877968

Observation 850e99cd-29f6-4007-a1c9-33882eea3cf1 · outbound

This paper cites and Ben-David, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Ben-David, S

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.934021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.934021Z digest=sha256:ae79e22cac75845591625d2f14ca66ee42be082130a0a00be98d0b3006dbadae

Observation 9afe5e2f-6267-448a-b52c-d34ce187313a · outbound

This paper cites Mathematical Statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Mathematical Statistics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.022585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.022585Z digest=sha256:ef83ba6dab02b65641ebc7674287cfdb18306509ade1e044ba8de604e5a4cdad

Observation a5a578f7-f424-47bd-9594-7bb5521b9d41 · outbound

This paper cites On methods of sieves and penalization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation On methods of sieves and penalization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.002895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.081016Z digest=sha256:783d1c9ffe94138b37fb09b65f180bb46b03c0fb195548bda6208097c63b832e

Observation 067892b6-6dcf-460f-93c9-2e9dfe291a08 · outbound

This paper cites Deeply-debiased off-policy interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Deeply-debiased off-policy interval estimation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.855972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.143606Z digest=sha256:0f8778b3ee4b7b3d4305001a51429935c84faff41652dab9628f4e4dc3095444

Observation 3587ffd9-3294-4bd0-b23e-85ae9726279a · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.668572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.211350Z digest=sha256:34349f517a5620d593a4f3c10c2f84f6f88539bd110531b06808bfd1d2415562

Observation 337a06a4-da2c-4071-a7f1-ce0b1005ff5f · outbound

This paper cites Statistical inference of the value function for reinforcement learning in infinite-horizon settings.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Statistical inference of the value function for reinforcement learning in infinite-horizon settings

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.489721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.281999Z digest=sha256:79dd9fe924db98d8d047f44f45f665dbcd4533019b418b6c3d409f0942856825

Observation d688c1de-2597-498c-a6cf-596f4ba2351b · outbound

This paper cites Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.222749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.378112Z digest=sha256:dd9213060d347ebcb7a46dfe67c2e90ef2e3ea3f23e79d6ebcf39359081c8c5c

Observation d1a16f9e-e59f-46b5-a50d-514496496e92 · outbound

This paper cites Off-policy confidence interval estimation with confounded markov decision process.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy confidence interval estimation with confounded markov decision process

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.987934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.470043Z digest=sha256:1946487b24140f73b0e4f64a3d9a1588d7a3ca8f9a105ec5f16c0e0b864c5b9a

Observation 61258582-d81d-4811-8a87-20401a18103d · outbound

This paper cites ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.476262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.476262Z digest=sha256:65bbb4b71ebe93c6c288d405e9966aa9a9095876e95d971a812b6f2c118308fe

Observation 02bd134f-5140-496e-867d-2a6dc4b7aea0 · outbound

This paper cites S., Szepesv \'a ri, C., and Maei, H.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Szepesv \'a ri, C., and Maei, H

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.811401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.553538Z digest=sha256:395801fdad507ffa5d608ee05569b3f1a8fa35a4866218ed27ec70cdb81a6e46

Observation a266e348-4d60-4a95-a6c3-91f20e10358b · outbound

This paper cites Doubly robust bias reduction in infinite horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly robust bias reduction in infinite horizon off-policy estimation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.583530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.648961Z digest=sha256:34818953dad599c52e43ca9148d5072ee8c886867d7de698f346bb3e37a82352

Observation bd7a0f40-a240-49ca-9b93-4c4ed5bba7b4 · outbound

This paper cites Off-policy evaluation in partially observable environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in partially observable environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.354914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.740428Z digest=sha256:56654285aa25cd8d85ef782eadf375af43fa61eee7cf8f4cb3afc669551e69e4

Observation 6b8460eb-5193-457b-a1b5-abfd2a007cdb · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:08.090302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.839634Z digest=sha256:148ba64eb6ebc13624a9421c05b78c5582fe49cc8fa1e28611a2ebc769952218

Observation 01ac0eea-2e90-47dc-b9e5-108682734923 · outbound

This paper cites S., Theocharous, G., and Ghavamzadeh, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Theocharous, G., and Ghavamzadeh, M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.914057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.954745Z digest=sha256:801e2f453bf95be6c794ff4abf6b2631b96526108924a29f687d87b444722c6d

Observation 1319aa44-ad66-41b0-a89e-f24c72174927 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:07.764164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.070422Z digest=sha256:c838841842fb95fa677472d426718ec65feb57accea816a8268c2756bdd976b1

Observation 95427da5-953e-4828-a1a3-9afece42be71 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Minimax weight and q-function learning for off-policy evaluation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.582206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.160195Z digest=sha256:531c6339b51a8152aebf7dcfadb0a08ff736fe511e410c74c983d5ad41095e6c

Observation 545f29b9-a6c7-4efa-b8dc-a8ed35540f44 · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:59.244313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:59.244313Z digest=sha256:17e4ddd41c94bbf7bcc30f95414201377a5fadd8b9d61b7956cab30aa4760448

Observation 4faaad34-d831-47bf-9316-24161d54223a · outbound

This paper cites Future-dependent value-based off-policy evaluation in pomdps.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Future-dependent value-based off-policy evaluation in pomdps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.402734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.334783Z digest=sha256:efa254b00136656464ebeeb4ae621c6103234cb2d8f8f5bf2013f3ec340b133a

Observation dbca4510-59cb-4e97-9bcb-3d4bd89e501e · outbound

This paper cites W., Wellner, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., Wellner, J

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.227633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.448099Z digest=sha256:8ad8f49a463ceccc1ced24361545acd5c9f4f917dcd08c2b223c83f410900b96

Observation bce7b07e-cbc8-4d16-b354-076b5e4193d5 · outbound

This paper cites Safe exploration for efficient policy evaluation and comparison.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Safe exploration for efficient policy evaluation and comparison

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.029178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.535998Z digest=sha256:320279f6024e2351ea72fc1344d551f0c16849f4cec11eaa955ea8c514ecde09

Observation a245a698-85c7-4393-ba08-1b02e7f94aea · outbound

This paper cites Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.769580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.636730Z digest=sha256:c02f6648ac0dfc79bd4cc8b95787832e19633f1ec5985da2070307429f3e0b22

Observation 323967c1-c2f2-4265-87d9-3bc98a27f1c8 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:06.847343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.714007Z digest=sha256:7e317f9c075a574416b2b9f733b0f31cad4e58f230ffeb4ffff921cf4c49e159

Observation e4fc61c2-0fef-43a3-9468-12ebcf38ec60 · outbound

This paper cites Off-policy evaluation for tabular reinforcement learning with synthetic trajectories.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation for tabular reinforcement learning with synthetic trajectories

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.682959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.793756Z digest=sha256:5e9f918e58250248c4c18a2201662553b935d2e5fed734fd856c604c676a574b

Observation 563ceaa3-3eff-4226-94b7-4f6ddc600d0b · outbound

This paper cites Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.487321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.899396Z digest=sha256:bdb8d0d1108be3eadaf593d38ae6f6d131632d9edc19d8fbf9725cc9dadfe129

Observation babb5910-858b-4ed3-8b50-cc5542ed781b · outbound

This paper cites Semiparametrically efficient off-policy evaluation in linear markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Semiparametrically efficient off-policy evaluation in linear markov decision processes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.298936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.966690Z digest=sha256:c46293a4a8ac7c850d13f52498d6bade1309541b040ef74c9335864ba5e05a19

Observation ffa3eb27-0ac8-42a5-a7a2-7949cdab0a73 · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.117651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.089605Z digest=sha256:0dada5f58567d3f00587e28c2c1c4f3d5c1365683221ddc1008ab0748f245df3

Observation 748abfb2-46a1-4789-b90b-a28ee93d0e2f · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.956955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.188458Z digest=sha256:a97437c5ece7b02c2ad071c0480d68b2c21e6515844357303e179cd51600e3f9

Observation 70ddb5a8-7bf2-4b41-83e2-ba540f4ed7b1 · outbound

This paper cites Quantile Off-Policy Evaluation via Deep Conditional Generative Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Quantile Off-Policy Evaluation via Deep Conditional Generative Learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.401023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.283152Z digest=sha256:fd37ed49b8806840391d67b197f5fab4861e5bd45e134a16de322968bff21655

Observation 6de15f68-f83f-4b87-b49d-e48db67e2cf9 · outbound

This paper cites An instrumental variable approach to confounded off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation An instrumental variable approach to confounded off-policy evaluation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:00.380335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:00.380335Z digest=sha256:9eab8742a9f64bfa303e6f6df9d6c698b028a50a7029fe4ad19c6a802d07c704

Observation 84acc115-2bbd-4e54-8447-7ca1fb3ddd45 · outbound

This paper cites and Wang, Y.-X.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wang, Y.-X

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.792318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.494044Z digest=sha256:bb30b4f33f0db7035467e2a9830b06dbeefb483f456d503c766faa3e0033bbba

Observation e0498785-3fb3-4552-984a-f41cacf3261e · outbound

This paper cites Two-way deconfounder for off-policy evaluation in causal reinforcement learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Two-way deconfounder for off-policy evaluation in causal reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.635897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.581039Z digest=sha256:e517c7ccd336fe68eed51d60638005c865a3dc64476e9be9bafe53eaa220d2f0

Observation 9856e75b-2403-45c1-9963-e03bb168f867 · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.492543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.642581Z digest=sha256:eefad28a5dc04e6f7c2017a1111b682149a7cb6baff89d031953fddfb0bc5b79

Observation aff6e40f-7662-485f-bf4f-6aecd7d41bad · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.360563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.743625Z digest=sha256:a5ff1d3c0181ecf6893e3dabd259dc631cf520e411076821d6d9805b7c0f6c2a

Observation 091369b4-6ed4-49e7-9e1b-bbe140f8acdd · outbound

This paper cites and Zhang, Y.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, Y

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.217784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.837982Z digest=sha256:537981e069733af9b7c79f34dcf0493b90af24bfb0ba807191aac8761b534556

Observation 817f2570-b3f1-411c-8418-e5bb063fee49 · outbound

This paper cites B., and Kosorok, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation B., and Kosorok, M

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.066644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.918164Z digest=sha256:8e73cf7ad61eb4a7f04101e3bf88e328a38d8a5f19d041b3648c1575157a3fd0

Observation de491975-afd8-4969-9253-61c05ba1bdab · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:01.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.981742Z digest=sha256:066d0bc3fd5ef8a34e917d36d17b67ace8c90489f09dbcc3ffba21f210a537be

Observation 8c1e2da9-074b-421d-9278-2c4e3909cf93 · outbound

This paper cites Robust offline reinforcement learning with heavy-tailed rewards.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Robust offline reinforcement learning with heavy-tailed rewards

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:04.924461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:17:01.051859Z digest=sha256:c4aa06e88229e921f87805bf0966f17b1aafbf6cdd8462ad41df255a3b40b700

Observation 34152c82-94af-4f4e-a037-fa18debd6725 · outbound

This paper cites write newline.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation write newline

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:01.124208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:01.124208Z digest=sha256:29bb3de571af2b5f3acf2dd4af794e5c0f3c52051af61d720b2d9d8ebdf11835

Pith citing papers

No inbound Pith citation observations are available.