Pith. sign in

REVIEW 1 cited by

Regret Minimization and Statistical Inference in Online Decision Making with High-dimensional Covariates

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.06329 v2 pith:4J267LFF submitted 2024-11-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords inferencebanditregretalgorithmconditionconsistentdecision-makingestimator
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper investigates regret minimization, statistical inference, and their interplay in high-dimensional online decision-making based on the sparse linear context bandit model. We integrate the $\varepsilon$-greedy bandit algorithm for decision-making with a hard thresholding algorithm for estimating sparse bandit parameters and introduce an inference framework based on a debiasing method using inverse propensity weighting. Under a margin condition, our method achieves either $O(T^{1/2})$ regret or classical $O(T^{1/2})$-consistent inference, indicating an unavoidable trade-off between exploration and exploitation. If a diverse covariate condition holds, we demonstrate that a pure-greedy bandit algorithm, i.e., exploration-free, combined with a debiased estimator based on average weighting can simultaneously achieve optimal $O(\log T)$ regret and $O(T^{1/2})$-consistent inference. We also show that a simple sample mean estimator can provide valid inference for the optimal policy's value. Numerical simulations and experiments on Warfarin dosing data validate the effectiveness of our methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Experimental Design With Estimation-Regret Trade-off Under Network Interference

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A Pareto-optimal trade-off between regret and treatment-effect estimation is derived and achieved for bandits with network interference, by compressing the action space through exposure mapping.

Pith tools