Pith. sign in

REVIEW 1 cited by

A Simple and Provably Efficient Algorithm for Asynchronous Federated Contextual Linear Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.03106 v1 pith:TPXZPJWJ submitted 2022-07-07 cs.LG stat.ML

classification cs.LGstat.ML
keywords contextualcommunicationlinearagentsalgorithmasynchronousbanditsfederated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study federated contextual linear bandits, where $M$ agents cooperate with each other to solve a global contextual linear bandit problem with the help of a central server. We consider the asynchronous setting, where all agents work independently and the communication between one agent and the server will not trigger other agents' communication. We propose a simple algorithm named \texttt{FedLinUCB} based on the principle of optimism. We prove that the regret of \texttt{FedLinUCB} is bounded by $\tilde{O}(d\sqrt{\sum_{m=1}^M T_m})$ and the communication complexity is $\tilde{O}(dM^2)$, where $d$ is the dimension of the contextual vector and $T_m$ is the total number of interactions with the environment by $m$-th agent. To the best of our knowledge, this is the first provably efficient algorithm that allows fully asynchronous communication for federated contextual linear bandits, while achieving the same regret guarantee as in the single-agent setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Experimental Design With Estimation-Regret Trade-off Under Network Interference

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A Pareto-optimal trade-off between regret and treatment-effect estimation is derived and achieved for bandits with network interference, by compressing the action space through exposure mapping.

Pith tools