REVIEW 2 cited by
Cross-Validated Off-Policy Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study estimator selection and hyper-parameter tuning in off-policy evaluation. Although cross-validation is the most popular method for model selection in supervised learning, off-policy evaluation relies mostly on theory, which provides only limited guidance to practitioners. We show how to use cross-validation for off-policy evaluation. This challenges a popular belief that cross-validation in off-policy evaluation is not feasible. We evaluate our method empirically and show that it addresses a variety of use cases.
Forward citations
Cited by 2 Pith papers
-
Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits
COPE/COPE-PG, a cross-domain off-policy evaluation and learning method, leverages source-domain data to estimate and optimize target-domain policies even with few-shot data, deterministic logging, and completely new actions.
-
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
A new importance-weighted estimator, OPFV, estimates and optimizes future policy value in non-stationary bandit environments by leveraging recurring time features in historical logs.
Discussion (0). Continue with ORCID to comment.