Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2212.06355.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:38.664480Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
13
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation f38d36e8-384d-43cb-97c2-7292f3c6793c · inbound
Selective Reviews of Bandit Problems in AI via a Statistical View A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cd3a95-016f-47d2-9b2e-0b0a0477d58e · inbound
Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a821d29-24b8-4596-b473-292c7a7b348c · inbound
Decoupled Functional Central Limit Theorems for Two-Time-Scale Stochastic Approximation A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed861f4-25aa-4860-974f-4099b90c232b · inbound
A Graphical Approach to State Variable Selection in Off-policy Learning A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 688
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc4a5bf-7afe-4a1a-8fbe-e4c2880169d5 · inbound
Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6dff4d6-33d6-4d14-98ef-18243337a3db · inbound
Just Trial Once: Ongoing Causal Validation of Machine Learning Models A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdbf2c37-bfe9-4204-9af2-6cf37b9050df · inbound
A Review of Causal Decision Making A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01ee2d5f-1c58-4ae7-bb69-e5c3ef572370 · inbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b592663-8483-42d3-941e-c08e82b2e266 · inbound
Q-function Decomposition with Intervention Semantics with Factored Action Spaces A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ebb3824-b34c-4ddb-a2ba-ceddff9ba563 · inbound
Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 9668
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623338d5-8695-4c94-bd4c-57012b7e388d · inbound
STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545f29b9-a6c7-4efa-b8dc-a8ed35540f44 · inbound
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0975775b-e2e8-46ae-a211-5ebf3d29d9fe · inbound
Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec131fbf-4ca2-419c-bc19-dbc1f811ceb2 · inbound
A General Framework for Off-Policy Learning with Partially-Observed Reward A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5126c5e8-f0fc-41d3-8d70-f7a54f035b15 · inbound
Dilution, Diffusion and Symbiosis in Spatial Prisoner's Dilemma with Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71f77b42-b88d-45cb-9266-814c6c497742 · inbound
Off-Policy Evaluation and Learning for Matching Markets A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf10217-2875-4625-a158-f035ca1bd6db · inbound
PAC Off-Policy Prediction of Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a21733a-6d48-4443-b10c-18d2dea51e21 · inbound
A Two-armed Bandit Framework for A/B Testing A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb417b0-945c-4f23-a185-22671bba7419 · inbound
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21904c4c-9910-44da-bd85-a2836baf8866 · inbound
Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6dcb46-8b48-4b56-8d09-4d526e5c7b29 · inbound
Hadronic screening masses in thermal QCD up to the electroweak scale A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027dc780-29f4-45c1-a275-1c9f45d70988 · inbound
Off-Policy Learning with Limited Supply A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8022a691-20d7-4621-bc51-7288bba4170f · inbound
Off-Policy Learning with Limited Supply A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9e757676-350e-496d-b3ca-cd690c2cf4c0 · inbound
Distributional Off-Policy Evaluation with Deep Quantile Process Regression A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a26eaaae-39f6-4f11-97c5-4b08e03d2779 · inbound
An adaptive variance estimator for relative sparsity A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 241
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c449f5df-6fa4-4c3e-bef6-817e81e6fd42 · inbound
Logging Policy Design for Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50ff0ba2-8551-46bf-a797-7fbf8e18210a · inbound
Logging Policy Design for Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 624cbe2f-a7d7-404e-9bda-295f8c7a7e4c · inbound
Offline Contextual Bandits in the Presence of New Actions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f40bba13-9c9b-4dfb-a913-204a65a5d85b · inbound
Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 19a8af3d-0ca2-4ce8-919d-1d09008e38b1 · inbound
Precision Physical Activity Prescription via Reinforcement Learning for Functional Actions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0aca948-a556-4307-a27e-022f73be3e7c · inbound
Counterfactually Safe Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5867cd6f-775f-49c8-976a-29c92c3f5270 · inbound
Bandit Simulation for Average Reward Inference A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation be0f2cab-3df7-4ccd-bea7-2b89e66b6a8b · inbound
Off-Policy Evaluation with Strategic Agents via Local Disclosure A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d9b28a7f-849b-4fd2-9ba8-121c30bbeba1 · inbound
Anytime-valid Optimal Policy Identification A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 39e8d999-992d-4257-866e-1528db3ffa24 · inbound
Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b769942f-d1d5-4cdd-863a-18578ec8467c · inbound
Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8fe23c9-8a2e-41a7-8ce8-dd776dc4186f · inbound
Fitted Occupancy-Ratio Evaluation without Bellman Completeness A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 284b73ce-0129-48df-964e-0511d0912fb1 · inbound
Fitted Occupancy-Ratio Evaluation without Bellman Completeness A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6315ae6e-de2c-4731-97dc-b49e5b50e184 · inbound
Estimating Causal Effects from Data Generated by Stochastic Algorithms A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 491ff237-264e-4e00-a65b-322be70aa7fa · inbound
A Statistical Test for the Benefits of Personalizing Interventions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43878d4-3aa2-42f8-9c17-ec9da84d492d · inbound
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3aa8a5-109a-4a65-81eb-688dc7111db7 · inbound
Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce8b195-a019-46e1-a093-05fdf0f79678 · inbound
Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.