Pith. sign in

REVIEW 6 cited by

A Tutorial on Thompson Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1707.02038 v3 pith:P4UWVK35 submitted 2017-07-07 cs.LG

classification cs.LG
keywords problemsalgorithminformationsamplingthompsonactionsdecisionlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Thompson sampling is an algorithm for online decision problems where actions are taken sequentially in a manner that must balance between exploiting what is known to maximize immediate performance and investing to accumulate new information that may improve future performance. The algorithm addresses a broad range of problems in a computationally efficient manner and is therefore enjoying wide use. This tutorial covers the algorithm and its application, illustrating concepts through a range of examples, including Bernoulli bandit problems, shortest path problems, product recommendation, assortment, active learning with neural networks, and reinforcement learning in Markov decision processes. Most of these problems involve complex information structures, where information revealed by taking an action informs beliefs about other actions. We will also discuss when and why Thompson sampling is or is not effective and relations to alternative algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Counterfactual Explanation of Shapley Value in Data Coalitions

    cs.GT 2025-07 conditional novelty 6.0 of 10

    The paper defines counterfactual explanations for Shapley values in data coalitions and proposes SV-Exp, a greedy algorithm that efficiently finds small data transfers to flip the value ranking.

  2. Learning Prosumer Behavior in Energy Communities: Integrating Bilevel Programming and Online Learning

    math.OC 2025-01 conditional novelty 6.0 of 10

    A bilevel price-setting optimizer is integrated with Thompson sampling, allowing an energy community manager to learn individual prosumer signature weights from daily price responses without large historical datasets.

  3. Behaviour Suite for Reinforcement Learning

    cs.LG 2019-08 accept novelty 6.0 of 10

    bsuite is a set of diagnostic reinforcement learning experiments with automated scoring and analysis tools for core agent capabilities.

  4. On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...

  5. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    A VLA policy with auxiliary success/progress heads and AWR+RECAP-style RL finished 1st in the LeHome 2026 simulation round and 2nd on the real robot.

  6. Decision-Making Under Complete Uncertainty: You Will Regret Not Being Greedy

    cs.GT 2025-02 reject novelty 5.0 of 10

    The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, bu...

Pith tools