REVIEW 3 cited by
Near-Optimal Algorithms for Differentially Private Online Learning in a Stochastic Environment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
In this paper, we study differentially private online learning problems in a stochastic environment under both bandit and full information feedback. For differentially private stochastic bandits, we propose both UCB and Thompson Sampling-based algorithms that are anytime and achieve the optimal $O \left(\sum_{j: \Delta_j>0} \frac{\ln(T)}{\min \left\{\Delta_j, \epsilon \right\}} \right)$ instance-dependent regret bound, where $T$ is the finite learning horizon, $\Delta_j$ denotes the suboptimality gap between the optimal arm and a suboptimal arm $j$, and $\epsilon$ is the required privacy parameter. For the differentially private full information setting with stochastic rewards, we show an $\Omega \left(\frac{\ln(K)}{\min \left\{\Delta_{\min}, \epsilon \right\}} \right)$ instance-dependent regret lower bound and an $\Omega\left(\sqrt{T\ln(K)} + \frac{\ln(K)}{\epsilon}\right)$ minimax lower bound, where $K$ is the total number of actions and $\Delta_{\min}$ denotes the minimum suboptimality gap among all the suboptimal actions. For the same differentially private full information setting, we also present an $\epsilon$-differentially private algorithm whose instance-dependent regret and worst-case regret match our respective lower bounds up to an extra $\log(T)$ factor.
Forward citations
Cited by 3 Pith papers
-
Faster Rates for Private Adversarial Bandits
By batching losses and using heavy-tailed bandit algorithms, any non-private adversarial bandit algorithm can be made epsilon-differentially private with regret O(sqrt(KT)/sqrt(epsilon)), and the first private expert-...
-
Optimal Regret of Bernoulli Bandits under Global Differential Privacy
For epsilon-global-DP Bernoulli bandits, the paper proves a tighter lower bound and matching upper bounds (up to a factor alpha that can approach 1) using a new quantity d_epsilon and a new DP-Chernoff concentration i...
-
Connecting Thompson Sampling and UCB: Towards More Efficient Trade-offs Between Privacy and Regret
A new bandit algorithm, DP-TS-UCB, achieves a tunable privacy-regret trade-off, improving the privacy guarantee of Gaussian Thompson Sampling from O(sqrt(T)) to O(T^0.25) while preserving near-optimal regret.
Discussion (0). Continue with ORCID to comment.