Pith. sign in

REVIEW 3 cited by

Statistical Bootstrapping for Uncertainty Estimation in Off-Policy Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.13609 v1 pith:RVC7AEKP submitted 2020-07-27 cs.LG stat.ML

classification cs.LGstat.ML
keywords bootstrappingconditionsconfidenceintervalspolicystatisticalvalueyield
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In reinforcement learning, it is typical to use the empirically observed transitions and rewards to estimate the value of a policy via either model-based or Q-fitting approaches. Although straightforward, these techniques in general yield biased estimates of the true value of the policy. In this work, we investigate the potential for statistical bootstrapping to be used as a way to take these biased estimates and produce calibrated confidence intervals for the true value of the policy. We identify conditions - specifically, sufficient data size and sufficient coverage - under which statistical bootstrapping in this setting is guaranteed to yield correct confidence intervals. In practical situations, these conditions often do not hold, and so we discuss and propose mechanisms that can be employed to mitigate their effects. We evaluate our proposed method and show that it can yield accurate confidence intervals in a variety of conditions, including challenging continuous control environments and small data regimes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Model-based Bootstrap of Controlled Markov Chains

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    A model-based bootstrap for finite CMCs is distributionally consistent for transitions and, via the delta method, for OPE/OPR targets under nonstationary behavior policies.

  2. STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    STITCH-OPE uses stitched diffusion-generated sub-trajectories with negative behavior-policy guidance to perform off-policy evaluation in high-dimensional, long-horizon tasks.

  3. Trajectory World Models for Heterogeneous Environments

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A pre-trained trajectory world model with interleaved temporal and variate attention achieves positive transfer across heterogeneous control environments, improving transition prediction, off-policy evaluation, and mo...

Pith tools