Pith. sign in

REVIEW 1 cited by

Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03434 v1 pith:DZSQYMFM submitted 2024-06-05 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords importanceestimatorcommonframeworkpac-bayesianpessimismpolicygeneralization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Off-policy learning (OPL) often involves minimizing a risk estimator based on importance weighting to correct bias from the logging policy used to collect data. However, this method can produce an estimator with a high variance. A common solution is to regularize the importance weights and learn the policy by minimizing an estimator with penalties derived from generalization bounds specific to the estimator. This approach, known as pessimism, has gained recent attention but lacks a unified framework for analysis. To address this gap, we introduce a comprehensive PAC-Bayesian framework to examine pessimism with regularized importance weighting. We derive a tractable PAC-Bayesian generalization bound that universally applies to common importance weight regularizations, enabling their comparison within a single framework. Our empirical results challenge common understanding, demonstrating the effectiveness of standard IW regularization techniques.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Information-Theoretic Generative Clustering of Documents

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A clustering method that replaces document embeddings with language-model probabilities over generated texts achieves state-of-the-art results on four document datasets.

Pith tools