Pith. sign in

Paper Citation Record · LEDGER

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

As of 17 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2505.22492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22492 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:17:01.124208Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact7
  • verified fuzzy50
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe970c9f-0548-4569-aaa3-4eab2ff15df9 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:52.903422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:52.903422Z digest=sha256:cea6a63a578085f50e8dd97d1f1729759971a6705c172476c07b0817c105bf77

Observation b62aa258-e78b-4938-bb97-799ce524f0ef · outbound

This paper cites and Kallus, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Kallus, N

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.021013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.021013Z digest=sha256:bb9314eff6d0fe646646ee4d1d87238ab17449bf878f9f1b67b46bccac4d2320

Observation 3b741a11-b4dc-4569-b70d-945c5839b408 · outbound

This paper cites Off-policy evaluation in doubly inhomogeneous environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in doubly inhomogeneous environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.103572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.103572Z digest=sha256:202d12aab598f3d66585c2e493ad1bf5f75cada7abad687f4b0b1f3c9f26505a

Observation 440371c0-981b-4670-bec6-f0aacf55fdae · outbound

This paper cites More efficient off-policy evaluation through regularized targeted learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More efficient off-policy evaluation through regularized targeted learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.186263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.186263Z digest=sha256:fd8ab1c067ef408c7b7617d4f994a3ebac7d33f967d4789298a94b276ddd8bcc

Observation 342ab1ad-17f1-455a-9f63-4a19cc565efe · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.264406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.264406Z digest=sha256:312fb1c2d9df4707bb2ab834caf58e890acf96aa0358608d983185c62555ead7

Observation 552f657a-6eda-4337-a6d5-a218da5e2d98 · outbound

This paper cites OpenAI Gym.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation OpenAI Gym

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.388701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.388701Z digest=sha256:92e7c298b128b7129436bbf90aef0aeb08d043c9e3e6cda95a168d1d745d3ff9

Observation c6c5e743-b5b3-4cbb-aebf-ad81d8d95b29 · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:17:04.709883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.496289Z digest=sha256:95adc953d839311b780d72363223f3d146077c99c4af0568526efe1439d1f7f7

Observation 8a06d0b8-b9d7-40c3-916f-86b10aa7c50a · outbound

This paper cites Structured Difference-of-Q via Orthogonal Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Structured Difference-of-Q via Orthogonal Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.447238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.610687Z digest=sha256:08c4ae6807e00ecff40f43ab56ef543a22d3329187bd80ea2938e8b7a9f166df

Observation 67a52a5d-4608-4582-a6c2-bd0f88c6a259 · outbound

This paper cites and Berger, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Berger, R

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.686598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.686598Z digest=sha256:1fb88162295088c09266514e35c4e854f388e4969a2b4289157af4e5f4d84a48

Observation bc2ae633-0ad7-44ca-a0db-07a0658b0efa · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.794136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.794136Z digest=sha256:4e43c792005732eecccac6691ddfd64b84988f402efadc9710cfcdd7fa4b7510

Observation 49de3aaf-d1f8-4dbd-a4a4-759ad9b84b90 · outbound

This paper cites and Jiang, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Jiang, N

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.924369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.924369Z digest=sha256:11dfdc8bb352b509675035b0dbd5fc9f2f5d88f5873698c38b58b94371bdef08

Observation b8a8b783-825e-4208-95fc-5d209ea87241 · outbound

This paper cites and Qi, Z.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Qi, Z

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.023087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.023087Z digest=sha256:513e261b65a3d5bce6e94fe2fa5bf5e22ab6c0bde4e0a5e2071e06922424129f

Observation fca476a0-cfa0-4d92-ab2b-9bf5fb2ca2f7 · outbound

This paper cites Gaussian approximation of suprema of empirical processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Gaussian approximation of suprema of empirical processes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.133862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.133862Z digest=sha256:d0f319f4adb84a127be652102ab3262c3680d98f953bdccbb4fda61bca67cd58

Observation 52b4bb87-f570-4feb-84e1-1403a8ab136a · outbound

This paper cites Double/debiased machine learning for treatment and structural parameters.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Double/debiased machine learning for treatment and structural parameters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.246149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.246149Z digest=sha256:24736ee5afdd117ed7652c5b0c60e8c8019ecf3358b0a6a465691b7161cf6f0e

Observation db63a9a8-f309-47aa-acac-bd6bb2d1c379 · outbound

This paper cites Coindice: Off-policy confidence interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Coindice: Off-policy confidence interval estimation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.339217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.339217Z digest=sha256:2aafcd089971fc27175fbd1eefd9462c20f7ea1f71062521fd3f8b35ef7d1c9f

Observation 0a9c8274-3776-490b-a0ed-558411d6bf08 · outbound

This paper cites Doubly Robust Policy Evaluation and Optimization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Robust Policy Evaluation and Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.432077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.432077Z digest=sha256:07aa32c36ee1f7928b70ee3d728b5edef467cdd5c912f92127edf1891a759893

Observation 8446019f-3832-4df8-9d51-6b1368b142b9 · outbound

This paper cites A theoretical analysis of deep q-learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A theoretical analysis of deep q-learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.524675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.524675Z digest=sha256:cd940f19aed8365f8d5864537ffc5a9d02b7a43eefe7ede615415660d8352c35

Observation 25caa724-aca7-4a07-b9f0-dda5c47cd091 · outbound

This paper cites More Robust Doubly Robust Off-policy Evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More Robust Doubly Robust Off-policy Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.617714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.617714Z digest=sha256:137912d2fa2561032d786bc2a1382bedc3d27c032cf003e24e3c5e1c7c5522c6

Observation 69d160b1-511b-4984-a4d4-ec8006edc122 · outbound

This paper cites Accountable off-policy evaluation with kernel B ellman statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Accountable off-policy evaluation with kernel B ellman statistics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.732514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.732514Z digest=sha256:34f47342b5baaca38ca27a5fcfe2e3819c5a1380ffbe504c7fe520910b83a200

Observation 1d2f3fc2-4499-45b8-ae5c-0061bd1e4779 · outbound

This paper cites Combining parametric and nonparametric models for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Combining parametric and nonparametric models for off-policy evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.828722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.828722Z digest=sha256:88657d730d2ef9ae8ae417b56d4ff0b7a13c9ea2594ae49822076ced2be46c1d

Observation 3e674bf6-aede-4609-a97b-8498caf0fb30 · outbound

This paper cites D., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation D., Thomas, P

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:15.124519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.910774Z digest=sha256:8177ce09bcb76b2952aedfd2a666a0fb1ed5ef8cedf8a3bb86b8bf36d885a0d0

Observation 238d6f10-29ef-4d5b-b3bf-b06331b78393 · outbound

This paper cites Importance sampling policy evaluation with an estimated behavior policy.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling policy evaluation with an estimated behavior policy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.980638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.995797Z digest=sha256:b0978b2885f21decb46c338d3ef4334a0a9a6a1b660d20f9ebb18ea994f0cdb7

Observation 715be3ee-ce98-4c84-ac12-c5babdbdda54 · outbound

This paper cites P., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Thomas, P

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.829766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.071207Z digest=sha256:49061da11bbc13535155bc586c6de5e4db063f4350dfa3d9309894ebbfe10d5a

Observation 5db9c47a-84a6-4898-9dc1-e9c7a05cdcd1 · outbound

This paper cites P., Niekum, S., and Stone, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Niekum, S., and Stone, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.679690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.196586Z digest=sha256:4aed47a9aa08e22c287ca9c139a76ae3197e11f33ebebf9adb1b9e171ea62ae9

Observation 9c90fb45-6232-4e37-b518-1c03af849c8c · outbound

This paper cites Bootstrapping fitted q-evaluation for off-policy inference.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Bootstrapping fitted q-evaluation for off-policy inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.525737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.317248Z digest=sha256:b027ab3d390dca06606bc0c0ecffd4f2c47cdc57bea4a6160ca3ed4a91ed916e

Observation 891b15a5-8bf1-45dc-abba-88a38f30a1b1 · outbound

This paper cites Importance sampling via the estimated sampler.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling via the estimated sampler

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.384125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.387170Z digest=sha256:30b445471f5f56097d78e9f3648176732cc0929a4e39f41280a6eb1bc242bdac

Observation 2ad6d719-9b29-446e-a5d2-f9e1a16fc9ca · outbound

This paper cites W., and Ridder, G.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., and Ridder, G

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.222790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.459117Z digest=sha256:0644e84a4e23c259e4b03ab2e953ee88dc22f2f4e366341aa214c579048f4f69

Observation db7f3592-4268-4450-ba42-a0663dfe4af2 · outbound

This paper cites and Wager, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wager, S

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:55.534857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:55.534857Z digest=sha256:b8024c07b2a97c41214a72ba942ca4d59536630d0feffe41737f605c2d51f177

Observation ff3d687f-0b80-4dd4-a4d7-db5b0d93866d · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.029859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.616369Z digest=sha256:113d362ab9806069ab68a2f4ec7f9a2ed0a4d4a021102f6266feb5e4135e9d8b

Observation a42c0dd8-5648-4354-b424-84d1a4e77faa · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.838362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.716241Z digest=sha256:73fa9888e1bea1164bc8c34623657f043610eede0190dc2ad8fefac6f4cd492b

Observation d404d600-7de2-4150-946b-5307bac95f8c · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.626828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.793014Z digest=sha256:7901cc7787896f9946667b896543310591bd128bec937571865e90f967dddc16

Observation 6411d0f2-b200-46ad-b612-6dd15b17e52c · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.420849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.872168Z digest=sha256:c40d39c6eb260201f4e24a6a768a4e47952c4f948e6dea462a3fff5317a222e1

Observation 63221bff-3c62-45ab-9178-79969858e05b · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:13.160576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.932696Z digest=sha256:ddaa7a87cd3d106a6ce200b69493160b39d96915c79928aef1325802f37edbb2

Observation 5a9cc8b5-daf7-428d-9acd-b4d241a9daf0 · outbound

This paper cites Batch policy learning under constraints.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Batch policy learning under constraints

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.937701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.017006Z digest=sha256:4920cc0053c86800e28a109334a3858488d5b8e99ab9164324cd4c375b651acc

Observation 1924c1ed-bc7b-4f1b-98f6-816e702217d1 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.144780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.144780Z digest=sha256:079a2affa0a1b9f6fd58b1f36e2d95c265cbf0f541a6b44a7d7d77e6392a573e

Observation 977884f7-69c5-4e80-aa98-3b38ff7c87f1 · outbound

This paper cites High-probability sample complexities for policy evaluation with linear function approximation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation High-probability sample complexities for policy evaluation with linear function approximation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.262424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.248494Z digest=sha256:0622b5991dfe4f7037b030906d254d6104c8c40bdf53e07c59e7c0c9abe10299

Observation 38ccf486-2bf4-47ce-91d8-e2cb5ef3fff3 · outbound

This paper cites Off-policy estimation of long-term average outcomes with applications to mobile health.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy estimation of long-term average outcomes with applications to mobile health

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.773972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.359148Z digest=sha256:1075fd26e2f9e1fecfc45528f56557516a3f40acd19a43f6eea8d77624726a95

Observation 67f38c91-6c96-4360-9f51-f15cbc3b435a · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:12.624851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.466433Z digest=sha256:e104ad8721781d50caf43cfa5cdaef2852f5874f60efcb31020acc5c8898ce12

Observation 03b93b2d-2ad1-411f-9a43-8b252edd6bf6 · outbound

This paper cites Breaking the curse of horizon: infinite-horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Breaking the curse of horizon: infinite-horizon off-policy estimation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.478362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.565112Z digest=sha256:a8022e4220626ae8e7627f4245d9186e9e9ea1b05cee84a05d769419e1e584f5

Observation b1b84187-62f8-49de-bd51-6289711aa20c · outbound

This paper cites and Zhang, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.254767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.671738Z digest=sha256:ac7388621fa41ba0c430755bf1f4e76d1c7a63b5b68d9d7dd5bb9367b15034fd

Observation 20212785-c506-4558-a002-57494379b523 · outbound

This paper cites Doubly Optimal Policy Evaluation for Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Optimal Policy Evaluation for Reinforcement Learning

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:17:03.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.740396Z digest=sha256:febc41ecd61133df8b4077dc394922619f3ddce3cc3afeda5dee7ff64141b314

Observation b6029ff5-5b9f-48b3-9947-4d1d444ba384 · outbound

This paper cites Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:03.551774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.834090Z digest=sha256:4f4561ce79aa3f4f4f224d42a2237f39825086cc7a54faf7d782ff10004bba6c

Observation 1ed53d7c-3231-44c0-8f6b-0feeedbf69dc · outbound

This paper cites J., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation J., Laber, E

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.946901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.946901Z digest=sha256:00f7a92558b5b36e9cf6b2a27a1742a68fce2cee18ee34d7d59c7a9fac2e45e2

Observation 548001c1-8ac1-463d-adbe-e8cb79d621a4 · outbound

This paper cites P., and Nowak, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., and Nowak, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.077014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.007129Z digest=sha256:520d6d63e38476b49f2a984b45542ff860805ca20a11726c026426538f3336a6

Observation e038af15-6e2e-40e6-8d48-d63427dd26d3 · outbound

This paper cites A., van der Laan, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., van der Laan, M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.923588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.074159Z digest=sha256:6540f94e9a6e3e61ba61049fcb8a457de93c821ebc225ca3427141abe29837fb

Observation cb654943-6911-4027-b699-509c2785a743 · outbound

This paper cites Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.691191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.135333Z digest=sha256:9d4cd87b9e6303ce97ed5edf04f617ebd0073b6cd3af708a307baa8ace82ee64

Observation 3bbf4a0c-7e88-437c-ba2e-7be99f60802d · outbound

This paper cites A Spectral Approach to Off-Policy Evaluation for POMDPs.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Spectral Approach to Off-Policy Evaluation for POMDPs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.195687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.195687Z digest=sha256:8d9e963c2ca7f7c05fbf6e129c09e2373d7a45fde8b1fe8b1cc0a89b23b85f05

Observation 877ba37c-7916-41a1-bbda-3021e7f63f2b · outbound

This paper cites Off-policy policy evaluation for sequential decisions under unobserved confounding.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy policy evaluation for sequential decisions under unobserved confounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.540247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.246652Z digest=sha256:f1004294597766c1debe0899ce38b70b10d2ae9378bb3f77efb72a90c052145b

Observation fe049d9e-3dad-43a7-876b-39568be90b97 · outbound

This paper cites K., Hsieh, F., and Robins, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation K., Hsieh, F., and Robins, J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.289965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.316362Z digest=sha256:0f16815f4f55f1b8b978b05676e2fa4002659103f4528c2173e1db99f8290a42

Observation 220033ba-5dc0-445d-b8e8-855b2dd8aecf · outbound

This paper cites Training language models to follow instructions with human feedback.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Training language models to follow instructions with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.362862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.362862Z digest=sha256:a0912851fbed462f2c1c63d0c1d4d55e9368bfeb17372f258fac05fba864c2ba

Observation 7e730401-bf8b-41ff-a17e-a81d69418f46 · outbound

This paper cites S., and Singh, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., and Singh, S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.104301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.449604Z digest=sha256:dd6908027bd54ccfe7a739f67b3d26e263954eb8ec374f1382f93e2ce9923de4

Observation c5644559-bd83-4111-bc92-4c2224a768b3 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.519872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.519872Z digest=sha256:cfc8f46a60f91e15f14894dff3049b87ec38f363a3cce3e10f58e9527c3d22f7

Observation eb28c104-9602-4e69-8561-9b0715d8a30c · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:10.808173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.582139Z digest=sha256:ceadef9c9b61db029985870e7f1fbbcadfdb2da7db9558f11364fb0108fe1887

Observation 4695e73b-5847-4e3f-82b9-e682b16dc01b · outbound

This paper cites Conditional importance sampling for off-policy learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Conditional importance sampling for off-policy learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.692678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.632221Z digest=sha256:672bad4783715176f0049e3ad92e4f51416e1435faa5093aac280684e75bbade

Observation a5078da9-a7fa-4985-85d6-b91f6d93ef28 · outbound

This paper cites G., and Dabney, W.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation G., and Dabney, W

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.457993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.718028Z digest=sha256:717fd3733f615fc303c43bd4d0957d70b824ef31d851a0c5b4cf67d82ac6d04f

Observation b9e6c4a0-196f-4998-8a2a-b23b71887564 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Proximal Policy Optimization Algorithms

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.801762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.801762Z digest=sha256:cd0eae7c600c7bc6fce44b23b6e54bd1f82a8d89818d168ef19195b609e9613f

Observation 7e7a7233-70a3-4e87-b622-49670b846ed1 · outbound

This paper cites Estimating the dimension of a model.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Estimating the dimension of a model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.220070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.866226Z digest=sha256:2e6c8ff5151eaa316dcf08975c8a667df4763955b827b9ec2146929e79d48eec

Observation 850e99cd-29f6-4007-a1c9-33882eea3cf1 · outbound

This paper cites and Ben-David, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Ben-David, S

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.934021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.934021Z digest=sha256:ae79e22cac75845591625d2f14ca66ee42be082130a0a00be98d0b3006dbadae

Observation 9afe5e2f-6267-448a-b52c-d34ce187313a · outbound

This paper cites Mathematical Statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Mathematical Statistics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.022585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.022585Z digest=sha256:ef83ba6dab02b65641ebc7674287cfdb18306509ade1e044ba8de604e5a4cdad

Observation a5a578f7-f424-47bd-9594-7bb5521b9d41 · outbound

This paper cites On methods of sieves and penalization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation On methods of sieves and penalization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.002895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.081016Z digest=sha256:fa56b3bcc15573d1a315559deac84d8e6855ffad5847ab25031c7fd35ba1c66b

Observation 067892b6-6dcf-460f-93c9-2e9dfe291a08 · outbound

This paper cites Deeply-debiased off-policy interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Deeply-debiased off-policy interval estimation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.855972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.143606Z digest=sha256:b927cfdf9201003785332dc4852e1675f13c7976e08e74474c5dfd02baf1ef14

Observation 3587ffd9-3294-4bd0-b23e-85ae9726279a · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.668572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.211350Z digest=sha256:ca9a633a3cd61cc254ba4a5f06123afe46b02ae726c412052a51073bbeea8cff

Observation 337a06a4-da2c-4071-a7f1-ce0b1005ff5f · outbound

This paper cites Statistical inference of the value function for reinforcement learning in infinite-horizon settings.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Statistical inference of the value function for reinforcement learning in infinite-horizon settings

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.489721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.281999Z digest=sha256:01919b791026f548663f5772650262f255ed9875fa67e738662a1650b9dde215

Observation d688c1de-2597-498c-a6cf-596f4ba2351b · outbound

This paper cites Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.222749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.378112Z digest=sha256:b4669f2698dba16c90d83defba9328399deefc7fc0cbdea7940017a161f5d2f0

Observation d1a16f9e-e59f-46b5-a50d-514496496e92 · outbound

This paper cites Off-policy confidence interval estimation with confounded markov decision process.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy confidence interval estimation with confounded markov decision process

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.987934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.470043Z digest=sha256:106c1084cbcc974813e6220587215f2a30d273310f27ae2fc968472a9ed681e3

Observation 61258582-d81d-4811-8a87-20401a18103d · outbound

This paper cites ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.476262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.476262Z digest=sha256:65bbb4b71ebe93c6c288d405e9966aa9a9095876e95d971a812b6f2c118308fe

Observation 02bd134f-5140-496e-867d-2a6dc4b7aea0 · outbound

This paper cites S., Szepesv \'a ri, C., and Maei, H.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Szepesv \'a ri, C., and Maei, H

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.811401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.553538Z digest=sha256:09a6d5fbe16f143f6b80db2261269725c2e1c85a658c3e9ebd8ddf1e523fb268

Observation a266e348-4d60-4a95-a6c3-91f20e10358b · outbound

This paper cites Doubly robust bias reduction in infinite horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly robust bias reduction in infinite horizon off-policy estimation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.583530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.648961Z digest=sha256:0d970cc8e837d03d9de72effc0fe00f4098e137a0de4150df6ea10613b886a02

Observation bd7a0f40-a240-49ca-9b93-4c4ed5bba7b4 · outbound

This paper cites Off-policy evaluation in partially observable environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in partially observable environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.354914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.740428Z digest=sha256:67c0bdd971774137ec2f1231d0760415bafde40eb6cddc1d2d5c8a5096befa17

Observation 6b8460eb-5193-457b-a1b5-abfd2a007cdb · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:08.090302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.839634Z digest=sha256:95bd99ffa1a984df6a37fdec5e358ab247fe075c52311ba0df24e262289c4988

Observation 01ac0eea-2e90-47dc-b9e5-108682734923 · outbound

This paper cites S., Theocharous, G., and Ghavamzadeh, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Theocharous, G., and Ghavamzadeh, M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.914057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.954745Z digest=sha256:e8ee782094bee4f09febe5ff9755153758ef113a4a9085237811ed4d64dd55c1

Observation 1319aa44-ad66-41b0-a89e-f24c72174927 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:07.764164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.070422Z digest=sha256:b40817aa8a1bb04470e1c39047367bae3806d10574ba0d9dcb4666533cfb9fa9

Observation 95427da5-953e-4828-a1a3-9afece42be71 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Minimax weight and q-function learning for off-policy evaluation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.582206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.160195Z digest=sha256:d8197860b66871eb1f9d76a56d40a95243d29af0c46ad712e55c9db0425cf3ad

Observation 545f29b9-a6c7-4efa-b8dc-a8ed35540f44 · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:59.244313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:59.244313Z digest=sha256:17e4ddd41c94bbf7bcc30f95414201377a5fadd8b9d61b7956cab30aa4760448

Observation 4faaad34-d831-47bf-9316-24161d54223a · outbound

This paper cites Future-dependent value-based off-policy evaluation in pomdps.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Future-dependent value-based off-policy evaluation in pomdps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.402734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.334783Z digest=sha256:2a7581a5192798fc20db481e7b8fbb703f9996ba82d063ad3a2e5faa17b67282

Observation dbca4510-59cb-4e97-9bcb-3d4bd89e501e · outbound

This paper cites W., Wellner, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., Wellner, J

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.227633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.448099Z digest=sha256:9eb66632159d9d08e486bd1a77b7a44ef537ae32144db6adf0c47ccb3884ff27

Observation bce7b07e-cbc8-4d16-b354-076b5e4193d5 · outbound

This paper cites Safe exploration for efficient policy evaluation and comparison.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Safe exploration for efficient policy evaluation and comparison

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.029178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.535998Z digest=sha256:4719269ecf065439bf042c986eb7381e01b5d927fcad0bc5507e665a92fd2de6

Observation a245a698-85c7-4393-ba08-1b02e7f94aea · outbound

This paper cites Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.769580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.636730Z digest=sha256:92385f2edb68bec4cf7190f8531fa7809fd0ee5c5d2b1432fe1a7a2feb3c66f6

Observation 323967c1-c2f2-4265-87d9-3bc98a27f1c8 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:06.847343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.714007Z digest=sha256:ac8bf7668984c5bd364f3e23d88491cbb578a8e7a552b1d8e9ef73cbb41c9e3e

Observation e4fc61c2-0fef-43a3-9468-12ebcf38ec60 · outbound

This paper cites Off-policy evaluation for tabular reinforcement learning with synthetic trajectories.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation for tabular reinforcement learning with synthetic trajectories

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.682959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.793756Z digest=sha256:988f1e0e1cc7d84e1131b6beca2f973269c2a0d0d0c2bf75eb889a602d927d4d

Observation 563ceaa3-3eff-4226-94b7-4f6ddc600d0b · outbound

This paper cites Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.487321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.899396Z digest=sha256:8c1e40df785bf2452538eb775bd5987ef730e8a5b170e969255cd46d4807794b

Observation babb5910-858b-4ed3-8b50-cc5542ed781b · outbound

This paper cites Semiparametrically efficient off-policy evaluation in linear markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Semiparametrically efficient off-policy evaluation in linear markov decision processes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.298936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.966690Z digest=sha256:d49a46ce8f96092a989f5cf5b67078aba2aa67fe28ad7decc2a11856ddb29fcb

Observation ffa3eb27-0ac8-42a5-a7a2-7949cdab0a73 · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.117651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.089605Z digest=sha256:b1d93cb8c00115e0cb439e93e1a70fd34fa3bfd55ebfe2794cab3e3c16829c36

Observation 748abfb2-46a1-4789-b90b-a28ee93d0e2f · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.956955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.188458Z digest=sha256:10c725daa46dcef1ad49cc96f5b511df1cbec26e97e2cc5ef2d492a46fd1e690

Observation 70ddb5a8-7bf2-4b41-83e2-ba540f4ed7b1 · outbound

This paper cites Quantile Off-Policy Evaluation via Deep Conditional Generative Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Quantile Off-Policy Evaluation via Deep Conditional Generative Learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.401023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.283152Z digest=sha256:5ee5a186e11604f57d0dd29b734c78e55eed30ef455b182a2251c15d5bcc18f0

Observation 6de15f68-f83f-4b87-b49d-e48db67e2cf9 · outbound

This paper cites An instrumental variable approach to confounded off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation An instrumental variable approach to confounded off-policy evaluation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:00.380335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:00.380335Z digest=sha256:9eab8742a9f64bfa303e6f6df9d6c698b028a50a7029fe4ad19c6a802d07c704

Observation 84acc115-2bbd-4e54-8447-7ca1fb3ddd45 · outbound

This paper cites and Wang, Y.-X.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wang, Y.-X

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.792318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.494044Z digest=sha256:0858cb0415b00ff552981eb321f712c8965549d8a8b9388b5533d1e944d990e5

Observation e0498785-3fb3-4552-984a-f41cacf3261e · outbound

This paper cites Two-way deconfounder for off-policy evaluation in causal reinforcement learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Two-way deconfounder for off-policy evaluation in causal reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.635897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.581039Z digest=sha256:3cff234314fd46ff9213c4455ce8c9afff61e8fb725ae794348f0509310bd00f

Observation 9856e75b-2403-45c1-9963-e03bb168f867 · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.492543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.642581Z digest=sha256:995200cf82e48e5d7f7089ba6c622ac0728434092304a800929effd67620c7f7

Observation aff6e40f-7662-485f-bf4f-6aecd7d41bad · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.360563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.743625Z digest=sha256:ea2c46ed6aff988b179ea692a63f9a65f74e5cb2b5462c3658d62a46971ea3aa

Observation 091369b4-6ed4-49e7-9e1b-bbe140f8acdd · outbound

This paper cites and Zhang, Y.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, Y

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.217784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.837982Z digest=sha256:b014d65559693f5ed77e54c97c79ff8616de70af3bc239a21090a52735f894f2

Observation 817f2570-b3f1-411c-8418-e5bb063fee49 · outbound

This paper cites B., and Kosorok, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation B., and Kosorok, M

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.066644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.918164Z digest=sha256:63ed2824255300d633a96f02d289d544747f0be6aa5b4efd1e0bd7eee00ee0d6

Observation de491975-afd8-4969-9253-61c05ba1bdab · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:01.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.981742Z digest=sha256:d22beb61a3feefd5c92e069fa7b118a577a155e6c93a7e7226763c1fe89a1e75

Observation 8c1e2da9-074b-421d-9278-2c4e3909cf93 · outbound

This paper cites Robust offline reinforcement learning with heavy-tailed rewards.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Robust offline reinforcement learning with heavy-tailed rewards

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:04.924461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:17:01.051859Z digest=sha256:2f764bbc420a15fd81187c2f1052fd46e4dc1fef056cd7d5a129985f1d36d403

Observation 34152c82-94af-4f4e-a037-fa18debd6725 · outbound

This paper cites write newline.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation write newline

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:01.124208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:01.124208Z digest=sha256:29bb3de571af2b5f3acf2dd4af794e5c0f3c52051af61d720b2d9d8ebdf11835

Pith citing papers

No inbound Pith citation observations are available.