REVIEW 6 cited by
A Tutorial on Thompson Sampling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Thompson sampling is an algorithm for online decision problems where actions are taken sequentially in a manner that must balance between exploiting what is known to maximize immediate performance and investing to accumulate new information that may improve future performance. The algorithm addresses a broad range of problems in a computationally efficient manner and is therefore enjoying wide use. This tutorial covers the algorithm and its application, illustrating concepts through a range of examples, including Bernoulli bandit problems, shortest path problems, product recommendation, assortment, active learning with neural networks, and reinforcement learning in Markov decision processes. Most of these problems involve complex information structures, where information revealed by taking an action informs beliefs about other actions. We will also discuss when and why Thompson sampling is or is not effective and relations to alternative algorithms.
Forward citations
Cited by 6 Pith papers
-
Counterfactual Explanation of Shapley Value in Data Coalitions
The paper defines counterfactual explanations for Shapley values in data coalitions and proposes SV-Exp, a greedy algorithm that efficiently finds small data transfers to flip the value ranking.
-
Learning Prosumer Behavior in Energy Communities: Integrating Bilevel Programming and Online Learning
A bilevel price-setting optimizer is integrated with Thompson sampling, allowing an energy community manager to learn individual prosumer signature weights from daily price responses without large historical datasets.
-
Behaviour Suite for Reinforcement Learning
bsuite is a set of diagnostic reinforcement learning experiments with automated scoring and analysis tools for core agent capabilities.
-
On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization
Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...
-
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
A VLA policy with auxiliary success/progress heads and AWR+RECAP-style RL finished 1st in the LeHome 2026 simulation round and 2nd on the real robot.
-
Decision-Making Under Complete Uncertainty: You Will Regret Not Being Greedy
The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, bu...
Discussion (0). Continue with ORCID to comment.