Pith. sign in

REVIEW 1 cited by

Online learning in MDPs with side information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1406.6812 v1 pith:T2YBFNGH submitted 2014-06-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords informationsideapplicationslearningonlineregretaccountalgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study online learning of finite Markov decision process (MDP) problems when a side information vector is available. The problem is motivated by applications such as clinical trials, recommendation systems, etc. Such applications have an episodic structure, where each episode corresponds to a patient/customer. Our objective is to compete with the optimal dynamic policy that can take side information into account. We propose a computationally efficient algorithm and show that its regret is at most $O(\sqrt{T})$, where $T$ is the number of rounds. To best of our knowledge, this is the first regret bound for this setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Catoni Contextual Bandits are Robust to Heavy-tailed Rewards

    stat.ML 2025-02 conditional novelty 7.0 of 10

    Contextual bandits with general function approximation can achieve regret scaling with cumulative reward variance and only logarithmically with the reward range, using Catoni robust mean estimators, with a matching lo...

Pith tools