Pith. sign in

REVIEW 5 cited by

A Survey on Contextual Multi-armed Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1508.03326 v2 pith:5ET44QHF submitted 2015-08-13 cs.LG

classification cs.LG
keywords contextualsurveyadversarialalgorithmalgorithmsanalyzeassumptionbandit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this survey we cover a few stochastic and adversarial contextual bandit algorithms. We analyze each algorithm's assumption and regret bound.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles

    cs.LG 2026-07 conditional novelty 7.0 of 10

    An OCO algorithm with only O(√T) static regret, pluggable as a preconditioner selector, recovers the classical O(1/√T) stationarity rate on smooth stochastic nonconvex problems and the O(T^{-2/7}) rate on nonsmooth ones.

  2. Robust Aggregation of Calibrated Forecasts

    econ.TH 2026-06 unverdicted novelty 7.0 of 10

    Introduces a robust max-min benchmark for aggregating calibrated forecasts that is LP-tractable, dominates OIH, and is attained by online algorithms under forecast-only feedback.

  3. Latent Order Bandits

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Latent order bandits require only a known partial order on actions within each latent state rather than full reward distributions, enabling UCB and posterior-sampling algorithms with regret bounds that match or exceed...

  4. Identifiable Latent Bandits: Leveraging observational data for personalized decision-making

    cs.LG 2024-07 unverdicted novelty 6.0 of 10

    Identifiable latent bandits apply nonlinear ICA to observational data to recover representations sufficient for inferring optimal actions in new instances, shortening exploration time.

  5. AutoPilot: Learning to Steer High Speed Robust BFT

    cs.DC 2026-06 unverdicted novelty 5.0 of 10

    AutoPilot uses decentralized reinforcement learning to continuously adjust BFT protocol parameters online, achieving 49.8% lower end-to-end latency than static defaults in dynamic environments.

Pith tools