Pith. sign in

Paper Citation Record · LEDGER

A Review of Off-Policy Evaluation in Reinforcement Learning

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2212.06355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.06355 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:38.664480Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f38d36e8-384d-43cb-97c2-7292f3c6793c · inbound

Selective Reviews of Bandit Problems in AI via a Statistical View cites this paper.

Selective Reviews of Bandit Problems in AI via a Statistical View A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-11T23:49:45.621188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:49:45.621188Z digest=sha256:ab8585649ac59714b4283beee759cb9417e04ea33b63e21a572d389f809a617d

Observation 28cd3a95-016f-47d2-9b2e-0b0a0477d58e · inbound

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed cites this paper.

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:52:51.143142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:52:51.143142Z digest=sha256:2be202f17f64270387f880057525438a726427f805e382180a88dc69feb9417c

Observation 1a821d29-24b8-4596-b473-292c7a7b348c · inbound

Decoupled Functional Central Limit Theorems for Two-Time-Scale Stochastic Approximation cites this paper.

Decoupled Functional Central Limit Theorems for Two-Time-Scale Stochastic Approximation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T05:56:52.732140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:56:52.732140Z digest=sha256:a8f56df2da11304550d545ce19fa682b05a7ce1e76544ef4d5644474b90a049f

Observation 1ed861f4-25aa-4860-974f-4099b90c232b · inbound

A Graphical Approach to State Variable Selection in Off-policy Learning cites this paper.

A Graphical Approach to State Variable Selection in Off-policy Learning A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 688

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:44.705573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:49:44.705573Z digest=sha256:d3b6672049f68a2d8bd379c72d0aae884dbfbc35f4cfffc844ffa166280316ca

Observation dcc4a5bf-7afe-4a1a-8fbe-e4c2880169d5 · inbound

Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making cites this paper.

Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:45.033105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:13:45.033105Z digest=sha256:4d54c2ff5bd3a71ac307b67eb557710ad4de08fac7b7fa4f9c4a84e53b9ee25e

Observation f6dff4d6-33d6-4d14-98ef-18243337a3db · inbound

Just Trial Once: Ongoing Causal Validation of Machine Learning Models cites this paper.

Just Trial Once: Ongoing Causal Validation of Machine Learning Models A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T21:30:18.952146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:30:18.952146Z digest=sha256:88c207261415786d46b54a2783a65bc2f95e027a49b58b4e5e84e561c8970555

Observation fdbf2c37-bfe9-4204-9af2-6cf37b9050df · inbound

A Review of Causal Decision Making cites this paper.

A Review of Causal Decision Making A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:47:26.298835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T02:47:21.582667Z digest=sha256:c4a8ede5cf241c66bf4a7e675693bdd0fccbdbdbbe6cfe989c9719c50b871fa9

Observation 01ee2d5f-1c58-4ae7-bb69-e5c3ef572370 · inbound

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning cites this paper.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.664480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.664480Z digest=sha256:b4ccbb972b75ad573935c769145a98ec7fb0b36453356268561058dde2f33dd1

Observation 1b592663-8483-42d3-941e-c08e82b2e266 · inbound

Q-function Decomposition with Intervention Semantics with Factored Action Spaces cites this paper.

Q-function Decomposition with Intervention Semantics with Factored Action Spaces A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:16:16.592567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:16:16.592567Z digest=sha256:89896585a69bda17492de8f8e4eaf1ce6cc8ef486cdddc526d2147bd12fa71c7

Observation 6ebb3824-b34c-4ddb-a2ba-ceddff9ba563 · inbound

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness cites this paper.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 9668

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:55.385777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:55.385777Z digest=sha256:d4dfc727a1df1fa43decfac812c80e5e830202bb022cc150d3b50b259c2ed919

Observation 623338d5-8695-4c94-bd4c-57012b7e388d · inbound

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation cites this paper.

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:06.637025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:06.637025Z digest=sha256:ee803e4e4e8584fa3d219033043bf56a102cbdca07b5888bf4ce5bdb3980ec46

Observation 545f29b9-a6c7-4efa-b8dc-a8ed35540f44 · inbound

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation cites this paper.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:59.244313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:59.244313Z digest=sha256:e177285a8b8774f9ccabea88d2fc9f8d7c8737e5f947a946c1ba43f9042d8440

Observation 0975775b-e2e8-46ae-a211-5ebf3d29d9fe · inbound

Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator cites this paper.

Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:36.368463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:30:36.368463Z digest=sha256:b041531a1194fae171e85e9de2a3605a043d3fd7ae11c79071c1160918933f1e

Observation ec131fbf-4ca2-419c-bc19-dbc1f811ceb2 · inbound

A General Framework for Off-Policy Learning with Partially-Observed Reward cites this paper.

A General Framework for Off-Policy Learning with Partially-Observed Reward A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.337409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:30.337409Z digest=sha256:e8da295b671fa12bdbab6be13f0ebfc246facf28be3a81a3dd8129eea2c85c9f

Observation 5126c5e8-f0fc-41d3-8d70-f7a54f035b15 · inbound

Dilution, Diffusion and Symbiosis in Spatial Prisoner's Dilemma with Reinforcement Learning cites this paper.

Dilution, Diffusion and Symbiosis in Spatial Prisoner's Dilemma with Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:06.764758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:06.764758Z digest=sha256:3a2f989e9ba708827d71234b2e125c89f5d4ff3c441cb983887cff6c724c52c7

Observation 71f77b42-b88d-45cb-9266-814c6c497742 · inbound

Off-Policy Evaluation and Learning for Matching Markets cites this paper.

Off-Policy Evaluation and Learning for Matching Markets A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:52.147480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:28:52.147480Z digest=sha256:4957f730eb4ed7185ee10fcc74f77095059e3777cdb0de8464eb740f3c35d838

Observation 3bf10217-2875-4625-a158-f035ca1bd6db · inbound

PAC Off-Policy Prediction of Contextual Bandits cites this paper.

PAC Off-Policy Prediction of Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:23:33.504561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:23:33.504561Z digest=sha256:2a0054f492d01adc6f78ab6e7be9a8d71d91af7838d12b6a7f373c1fe244cae7

Observation 7a21733a-6d48-4443-b10c-18d2dea51e21 · inbound

A Two-armed Bandit Framework for A/B Testing cites this paper.

A Two-armed Bandit Framework for A/B Testing A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:25.679066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:47:25.679066Z digest=sha256:9889e4271a5dcea51453a6d8c8b9db0e9bcee2faf0b9694515c216443ccdee64

Observation fcb417b0-945c-4f23-a185-22671bba7419 · inbound

GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents cites this paper.

GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:30:42.106294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:30:42.106294Z digest=sha256:5ac48e26e99b0c75d35bcfba0a0ac27e5f21df41a9c0af94baadeff429f69407

Observation 21904c4c-9910-44da-bd85-a2836baf8866 · inbound

Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions cites this paper.

Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T09:09:53.522461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:09:53.522461Z digest=sha256:c366533c0103b4e3dff0bf5e5386b0a6213e006200368176ccf75aa5e6bba203

Observation ca6dcb46-8b48-4b56-8d09-4d526e5c7b29 · inbound

Hadronic screening masses in thermal QCD up to the electroweak scale cites this paper.

Hadronic screening masses in thermal QCD up to the electroweak scale A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T22:27:27.987413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:27:27.987413Z digest=sha256:1db28159376d583b69b62d834d90ebf6c3ca42d587c40ec729451454613ebd39

Observation 027dc780-29f4-45c1-a275-1c9f45d70988 · inbound

Off-Policy Learning with Limited Supply cites this paper.

Off-Policy Learning with Limited Supply A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:19:52.629413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T08:16:29.736803Z digest=sha256:07f1b774f06512402226b9b6b020ebd1298982990f0f5787c6194ac4cf3266f3

Observation 8022a691-20d7-4621-bc51-7288bba4170f · inbound

Off-Policy Learning with Limited Supply cites this paper.

Off-Policy Learning with Limited Supply A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:20:02.388627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T11:17:13.734190Z digest=sha256:021bbdf9774dac22372202367318c2860b547cbf75d95ce8e8023f3139041c02

Observation 9e757676-350e-496d-b3ca-cd690c2cf4c0 · inbound

Distributional Off-Policy Evaluation with Deep Quantile Process Regression cites this paper.

Distributional Off-Policy Evaluation with Deep Quantile Process Regression A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:06:03.889981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T04:10:14.476158Z digest=sha256:447572c62e8f9042f95cbb7166363facb4cecdc507ca1cbec20395b29313dce5

Observation a26eaaae-39f6-4f11-97c5-4b08e03d2779 · inbound

An adaptive variance estimator for relative sparsity cites this paper.

An adaptive variance estimator for relative sparsity A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 241

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.544207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T19:35:46.097113Z digest=sha256:e8b173c8ca6ac77fa5628c042845c3e5446c48970fdf7695301f3cc12eb70df5

Observation c449f5df-6fa4-4c3e-bef6-817e81e6fd42 · inbound

Logging Policy Design for Off-Policy Evaluation cites this paper.

Logging Policy Design for Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:57.610094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T03:08:07.248759Z digest=sha256:da764da104ca7ce3cafc85f67c35258973a4543bcfa1d4ccb98358b9215d7bcc

Observation 50ff0ba2-8551-46bf-a797-7fbf8e18210a · inbound

Logging Policy Design for Off-Policy Evaluation cites this paper.

Logging Policy Design for Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:43:43.471489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T20:40:06.337842Z digest=sha256:a27abdec5a936e8ea560ac0d4bf2beed4500ff2f52a55cf0604ca3b9478666a2

Observation 624cbe2f-a7d7-404e-9bda-295f8c7a7e4c · inbound

Offline Contextual Bandits in the Presence of New Actions cites this paper.

Offline Contextual Bandits in the Presence of New Actions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.052928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:01:35.633356Z digest=sha256:476d7ba26df26ca400a528a02f66a8c42ec222800e2c68093b2abda9bc8e52cc

Observation f40bba13-9c9b-4dfb-a913-204a65a5d85b · inbound

Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions cites this paper.

Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T03:08:02.149098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T03:04:01.511997Z digest=sha256:bb5d8e665c4c2fdfcd86cee6047e44cc24e63b93d0788332d96aba434757febc

Observation 19a8af3d-0ca2-4ce8-919d-1d09008e38b1 · inbound

Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions cites this paper.

Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:47.766604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T18:16:37.532243Z digest=sha256:f4bd709b6ce5312c0b0897b08e27dc49b19f27d4307c32fdbfa2d5d1be2f6496

Observation c0aca948-a556-4307-a27e-022f73be3e7c · inbound

Counterfactually Safe Reinforcement Learning cites this paper.

Counterfactually Safe Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:04:06.866592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T23:48:41.383288Z digest=sha256:63648fc374227391eda9cc656af53739646865f4edd363d2fd31019723443a72

Observation 5867cd6f-775f-49c8-976a-29c92c3f5270 · inbound

Bandit Simulation for Average Reward Inference cites this paper.

Bandit Simulation for Average Reward Inference A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:27.352054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:50:08.211646Z digest=sha256:010aa3c80ca2b24b59c08d124643b9593ad829349801ca6f607e80320bb44417

Observation be0f2cab-3df7-4ccd-bea7-2b89e66b6a8b · inbound

Off-Policy Evaluation with Strategic Agents via Local Disclosure cites this paper.

Off-Policy Evaluation with Strategic Agents via Local Disclosure A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.409007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T22:02:00.194193Z digest=sha256:4097c58f2ea995f5998167bd848f34f57f539a1520a965a7d3844b751dc59994

Observation d9b28a7f-849b-4fd2-9ba8-121c30bbeba1 · inbound

Anytime-valid Optimal Policy Identification cites this paper.

Anytime-valid Optimal Policy Identification A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-26T23:50:15.173974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T23:49:05.019752Z digest=sha256:e18b56d14b5723e20e67f1432cc62661053862b7a18ee8c7d1635138b5c10502

Observation 39e8d999-992d-4257-866e-1528db3ffa24 · inbound

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random cites this paper.

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:39:40.721840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T15:38:16.806415Z digest=sha256:ee9395f2338c8ff59241564744d4e1d8dbc9ff06b810c6354cea577d277e1b75

Observation b769942f-d1d5-4cdd-863a-18578ec8467c · inbound

Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy Evaluation cites this paper.

Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:43.084562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T10:40:07.441340Z digest=sha256:08848a4885a1cec230732d7034908f43eb856e365a80a130cda7dd082c3b6f67

Observation d8fe23c9-8a2e-41a7-8ce8-dd776dc4186f · inbound

Fitted Occupancy-Ratio Evaluation without Bellman Completeness cites this paper.

Fitted Occupancy-Ratio Evaluation without Bellman Completeness A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:33:54.263949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-07T14:25:33.921937Z digest=sha256:ef9f96419c7beccd8b6e6e1cdaee0e5b989b03b342782d63bb619067f41a77a9

Observation 284b73ce-0129-48df-964e-0511d0912fb1 · inbound

Fitted Occupancy-Ratio Evaluation without Bellman Completeness cites this paper.

Fitted Occupancy-Ratio Evaluation without Bellman Completeness A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:33:36.437840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:33:36.437840Z digest=sha256:22f2fc64977d6b2882ce6124c9e0945189becb6e748646d4cb52ef88a1694f6b

Observation 6315ae6e-de2c-4731-97dc-b49e5b50e184 · inbound

Estimating Causal Effects from Data Generated by Stochastic Algorithms cites this paper.

Estimating Causal Effects from Data Generated by Stochastic Algorithms A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T00:15:47.443944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T00:15:30.364680Z digest=sha256:c747ba0d354cd3b07441fddaded4f30777316dbb68c2f2056ac8120c22054b25

Observation 491ff237-264e-4e00-a65b-322be70aa7fa · inbound

A Statistical Test for the Benefits of Personalizing Interventions cites this paper.

A Statistical Test for the Benefits of Personalizing Interventions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T05:33:39.540475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:33:39.540475Z digest=sha256:f37a0b76ca32aabebe7d52cb9b3e5ce0048ad699780d326d97a404aa35ab0cac

Observation f43878d4-3aa2-42f8-9c17-ec9da84d492d · inbound

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits cites this paper.

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:46.724676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:46.724676Z digest=sha256:0f29dcdc752cafc89a402e6a225e1caf10cb27f4ba1fa55c9b8d24f4bd2dde45

Observation fa3aa8a5-109a-4a65-81eb-688dc7111db7 · inbound

Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions cites this paper.

Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T03:07:12.773917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:07:12.773917Z digest=sha256:5ac7d1ba6b84a7ccbb088833ebacf74aa319088b07b1b291c7e76d7ad7950d03

Observation 1ce8b195-a019-46e1-a093-05fdf0f79678 · inbound

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions cites this paper.

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:15.568605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:15.568605Z digest=sha256:e3f38e15443b477d31684d303061989298ec3e43e40f6015086c3b75200e9929