REVIEW 3 major objections 5 minor 51 references
Multiplayer Bandit Learning, from Competition to Cooperation
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims that in a two-player one-armed bandit, competition lowers the exploration threshold below the solo Gittins index, cooperation raises it above, and neutral players can each beat the solo optimum by observing each other's…
desk verdict A useful unifying framework with mostly solid qualitative results, but Theorem 4 Part 2 has two real proof gaps and the almost-sure claims outrun the argument; worth refereeing, not yet citable as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are the Gittins index $g=g(\mu,\beta)$, the threshold known-arm probability at which a single player is indifferent between the known and risky arms; the copycat strategy, in which a player stays on the known arm until the opponent explores and then repeats the opponent's previous move one round later; and the induced threshold $p^*\leq (m\beta+g)/(1+\beta)$ where copying destroys the value of exploration. The cooperative result uses a lagged-copy arrangement that gives the team two observations per experiment. The neutral learning result is carried by a perfect Bayesian deviation argument: if a player received only the single-player optimum, she could switch, upon the opponent's first exploration, to the zero-sum copycat strategy and win in the remaining subgame, contradicting equilibrium. The long-term convergence results use a concentration inequality on the empirical mean of the risky arm to show that an infinitely exploring player eventually identifies the better arm and that the other player follows.
What would settle it
Compute, for a fixed prior $\mu$ and discount $\beta$, the value of the zero-sum game that starts with Alice forced to play the risky arm in round 0 and Bob forced to play the known arm. If for some $p\in(p^*,g)$ Bob's value is not strictly positive, the pivot of the neutral-learning theorem fails. A simulation counterpart is to truncate the game at a large horizon, solve for the perfect Bayesian equilibrium by dynamic programming, and check whether each player's expected reward still exceeds the single-player optimum on that interval.
Extended reading notes
Core claim
The paper's central claim is that the value of information in the two-player game is alignment-dependent. In the zero-sum regime ($\lambda=-1$) there is a threshold $p^*<g$, with the copycat bound $p^*\leq (m\beta+g)/(1+\beta)$, such that for every $p>p^*$ neither player ever explores in equilibrium; a player who experiments first hands the opponent a one-round-lagged copy of the information, and above $p^*$ that erases the explorer's advantage. Below the threshold $m+\beta w/2$, however, both players explore in the first round, so competition does not reduce play to pure myopia. In the fully cooperative regime ($\lambda=1$), one player can explore while the other copies with a delay, effectively turning one experiment into two observations, and this makes the team explore for some $p>g$. In the neutral regime ($\lambda=0$), for $p$ between the competing threshold $p^*$ and the solo Gittins index $g$, every perfect Bayesian equilibrium gives each player strictly more expected reward than a single player using an optimal strategy; the mechanism is that after the opponent's first exploration, a player can switch to a winning zero-sum strategy against her. Finally, in every Nash equilibrium competing and neutral players converge to the same arm with probability 1, whereas cooperating players have equilibria with infinitely many switches.
Load-bearing premise
The neutral-learning theorem rests on the unproved subgame property that after one player explores, the other can guarantee a strictly positive expected advantage for every $p$ in $(p^*,g)$; the paper's copycat argument proves this only for $p$ above the larger cutoff $(m\beta+g)/(1+\beta)$, so the interval between the two cutoffs is where the claim is unsupported.
Editorial extensions
If this is right
- Competing players will not explore the risky arm for any $p$ above $p^*$, so head-to-head rivalry can freeze experimentation even when a solo learner would continue.
- Competing players still explore for all sufficiently small $p$, so the zero-sum interaction does not collapse to always playing the safer arm.
- Two cooperating players can explore for values of $p$ where a single player would stop, because a lagged-copy arrangement makes one experiment yield two observations.
- Neutral players who observe actions but not rewards can each beat the single-player optimum in every perfect Bayesian equilibrium for $p\in(p^*,g)$, and the probability that no one explores decays exponentially.
- In every Nash equilibrium, competing and neutral players settle on the same arm almost surely, while cooperating players can have equilibria with one player switching arms infinitely often.
Reading between the lines
- If the copycat bound can be sharpened to cover the whole interval $(p^*,g)$, the neutral-learning theorem would pin the end of mutually beneficial learning exactly at the competition threshold; a direct numerical check is to compute the value of the zero-sum subgame after a forced exploration.
- The same copycat mechanism suggests that any rule raising the cost of imitation, such as a temporary protection for the arm a player first explored, would widen the exploration region for competing and neutral players.
- The exponential decay of the no-exploration probability gives a quantitative handle for algorithm designers: a finite-time agent can estimate equilibrium exploration probabilities and decide when observation has made further own-experimentation unnecessary.
- The non-convergence example for cooperating players indicates that aligned payoffs alone do not ensure coordinated specialization; equilibrium selection or communication would be needed to make teams settle on a single arm.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a two-player one-armed bandit in which, each round, a player chooses a predictable left arm (known success probability p) or a risky right arm (prior μ), observes own reward and the other player's action but not the other player's reward, and has utility Γ_i + λΓ_j. The three regimes are competing (λ=-1), neutral (λ=0), and cooperating (λ=1). The main claims are: competing players explore less than a single player for p above a threshold p* ≤ (mβ+g)/(1+β), yet still explore for p below about m+βw/2 (Theorems 1 and 2); cooperating players explore for some p above the Gittins index g (Theorem 3); neutral players explore with probability one for p<g and, in every perfect Bayesian equilibrium for p∈(p*,g), each player strictly beats the single-player optimum (Theorem 4); and competing and neutral players eventually settle on the same arm in every Nash equilibrium while cooperating players may oscillate (Theorems 9-10 and Proposition 2). Finite-horizon analogues and improved bounds for a uniform prior are also given.
Significance. If the proofs are completed, this is a valuable contribution to multiplayer bandit learning and strategic experimentation. The λ-interpolation gives a clean comparative framework, and the paper makes concrete falsifiable predictions: competition reduces exploration, cooperation increases it, neutral players can profit from observing each other's actions, and long-run agreement fails only outside the [-1,0] range. The paper also has genuine technical strengths: the copycat strategy is explicit, the concentration lemma (Lemma 4) is clean, and the thresholds are defined intrinsically rather than fitted to data. The main caveat is that the proof of the headline welfare result for neutral players, Theorem 4 Part 2, currently rests on two unjustified steps, and the algebra in Theorem 1 needs correction. These are fixable in principle, but they are load-bearing.
major comments (3)
- [Section 5, proof of Theorem 4 Part 2] The invocation 'By Theorem 1' does not cover all p∈(p*,g). Theorem 1's proof gives Bob a strict advantage via the copycat strategy only when p > (mβ+g)/(1+β), and the theorem only establishes p* ≤ (mβ+g)/(1+β). Thus the interval p∈(p*, (mβ+g)/(1+β)] is left unsupported. A separate argument is needed showing that in the zero-sum subgame where Alice is forced to play R at round k, Bob can guarantee a strictly positive net advantage for every p in (p*,g).
- [Section 5, proof of Theorem 4 Part 2] The assertion 'Since the equilibrium is perfect Bayesian, we have E(Γ'_A) ≥ α/(1−β)' does not follow. Alice's strategy S_A is a best response to S_B, not to the deviating strategy S'_B; under S'_B, which plays left until Alice explores, Alice may lose information that she exploited in the original equilibrium, so her payoff could fall below the single-player optimum α/(1−β). Without this lower bound, the inequality E(Γ'_B)>E(Γ'_A) does not imply that Bob's deviation beats α/(1−β), which is the contradiction the proof needs.
- [Section 3, Eqs. (7)-(8)] The algebraic identity leading to inequality (8) is incorrect. From the displayed expressions for E(Γ_A) and E(Γ_B) one obtains E(Γ_A)-E(Γ_B) = (m-p(1+β))β^k + (1-β)∑_{t=k+1}^∞ E(γ_A(t))β^t, not the displayed expression with an extra factor β^k on the tail sum. Consequently inequality (8) does not follow as stated. A corrected derivation changes the no-exploration threshold to (m+βg)/(1+β) (under the same bound on the tail), and Theorem 3's use of (8) inherits the problem. The stated theorems may still be true, but the proof must be redone.
minor comments (5)
- [Section 1.1 and Section 6.1] The model restricts λ to [-1,1], but Proposition 3 analyzes λ<-1; please clarify whether that proposition is intended as an out-of-model remark or whether the model should allow λ outside [-1,1].
- [Section 1.2.1, Theorem 1] The definition of p* as sup{p: arm R is explored in some Nash equilibrium} already makes 'for all p>p* the players do not explore' true by definition; the content of the theorem is the upper bound p*≤(mβ+g)/(1+β). The statement could be rephrased to avoid this redundancy.
- [Section 5, proof of Theorem 4 Part 1] The displayed chain from the equilibrium bound to the inequality Φ_k ≥ (α−p)/(1−p−β^k) is compressed; the term handling the case where exploration has already occurred is omitted in the first displayed inequality and only appears implicitly in the next line. Please spell out the derivation.
- [Section 6.1, Proposition 3] The sentence 'there is a perfect Bayesian equilibrium in which Bob visits both arms infinitely often whenever' ends abruptly; the trailing 'whenever' should be removed or completed.
- [References and typos] Several small typos remain: 'Rotschild' should be 'Rothschild' in Section 7, and 'Salomon' in the bibliography entry [RSV] should be 'Solan'. The reference [RSV] also lacks a year.
Circularity Check
No significant circularity; the central derivations are self-contained against the external Gittins-index benchmark.
full rationale
The paper's main claims are derived from the model definition, the Gittins index (an external benchmark), and equilibrium reasoning, rather than from fitted parameters or prior results by the authors. The thresholds p* and ~p are defined in terms of equilibrium exploration, but the nontrivial content, such as p* <= (m*beta + g)/(1 + beta) and ~p >= m + beta*w/2, is proved directly via copycat and value arguments; the no-exploration property above p* is a definitional consequence, not a fitted prediction. The neutral-player learning theorem (Theorem 4) is argued from the single-player Gittins optimum and a deviation argument, and the long-run convergence theorems use concentration and martingale-style bounds; none of these reduces to the statements being proved. Two steps in the proof of Theorem 4 Part 2 are not fully justified: the invocation of Theorem 1 for Bob's advantage in the range p in (p*, (m*beta + g)/(1 + beta)], and the assertion E(Gamma'_A) >= alpha/(1 - beta) under (S_A, S'_B) from perfect Bayesian rationality. These are correctness or rigor gaps, not circularity. The paper does not fit parameters to data, rename a known result, or import a uniqueness theorem from the authors' prior work.
Assumptions & free parameters
assumptions (5)
- domain assumption Gittins index theorem for the one-armed bandit, and the reduction of the single-player problem to comparing the risky arm's index g with the safe arm's success probability p.
- standard math Sion's minimax theorem guarantees the value of the zero-sum game.
- domain assumption Existence of perfect Bayesian equilibria in the neutral game (Fudenberg-Levine).
- domain assumption The information structure: players observe each other's actions but not rewards.
- domain assumption The prior mu has no atom at the safe arm's success probability p, i.e. mu(p) = 0.
Cite this review
Pith. "Pith review of Multiplayer Bandit Learning, from Competition to Cooperation." pith.science (2026). https://pith.science/paper/NT6QK3LV
@misc{pith2026190801135,
author = {Pith},
title = {Pith review of: Multiplayer Bandit Learning, from Competition to Cooperation},
year = {2026},
howpublished = {\url{https://pith.science/paper/NT6QK3LV}},
note = {Machine review of arXiv:1908.01135}
}
abstract
The stochastic multi-armed bandit model captures the tradeoff between exploration and exploitation. We study the effects of competition and cooperation on this tradeoff. Suppose there are $k$ arms and two players, Alice and Bob. In every round, each player pulls an arm, receives the resulting reward, and observes the choice of the other player but not their reward. Alice's utility is $\Gamma_A + \lambda \Gamma_B$ (and similarly for Bob), where $\Gamma_A$ is Alice's total reward and $\lambda \in [-1, 1]$ is a cooperation parameter. At $\lambda = -1$ the players are competing in a zero-sum game, at $\lambda = 1$, they are fully cooperating, and at $\lambda = 0$, they are neutral: each player's utility is their own reward. The model is related to the economics literature on strategic experimentation, where usually players observe each other's rewards. With discount factor $\beta$, the Gittins index reduces the one-player problem to the comparison between a risky arm, with a prior $\mu$, and a predictable arm, with success probability $p$. The value of $p$ where the player is indifferent between the arms is the Gittins index $g = g(\mu,\beta) > m$, where $m$ is the mean of the risky arm. We show that competing players explore less than a single player: there is $p^* \in (m, g)$ so that for all $p > p^*$, the players stay at the predictable arm. However, the players are not myopic: they still explore for some $p > m$. On the other hand, cooperating players explore more than a single player. We also show that neutral players learn from each other, receiving strictly higher total rewards than they would playing alone, for all $ p\in (p^*, g)$, where $p^*$ is the threshold from the competing case. Finally, we show that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while this can fail for cooperating players.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Multi-Player Bandits: The Adversarial Case
P. Alatur, K. Y. Levy, and A. Krause. Multi-player bandits: The adversarial case. arXiv preprint arXiv:1902.08036 , 2019
work page Pith review arXiv 1902
-
[2]
The perils of exploration under competition: A computational modeling approach
Guy Aridor, Kevin Liu, Aleksandrs Slivkins, and Zhiwei Steven Wu. The perils of exploration under competition: A computational modeling approach. In Proceedings of the 2019 ACM Conference on Economics and Computation, EC 2019, Phoenix, AZ, USA, June 24-28, 2019. , pages 171--172, 2019
work page 2019
-
[3]
Robert J. Aumann and Michael Maschler. Repeated Games with Incomplete Information . MIT Press, 1995
work page 1995
-
[4]
O. Avner and S. Mannor. Concurrent bandits and cognitive radio networks. In ECML/PKDD , 2014
work page 2014
-
[5]
M Anderson. The evolution of eusociality. Annual Review of Ecology and Systematics , 15(1):165--189, 1984
work page 1984
-
[6]
Mutual observability and the convergence of actions in a multi-person two-armed bandit model
Masaki Aoyagi. Mutual observability and the convergence of actions in a multi-person two-armed bandit model. Journal of Economic Theory , 82:405--424, 1998
work page 1998
-
[7]
Masaki Aoyagi. Corrigendum: Mutual observability and the convergence of actions in a multi-person two-armed bandit model. 2011
work page 2011
-
[8]
Toward a theory of discounted repeated games with imperfect monitoring
Dilip Abreu, David Pearce, and Ennio Stacchetti. Toward a theory of discounted repeated games with imperfect monitoring. Econometrica , 58(5):1041--1063, 1990
work page 1990
Show all 51 references
-
[9]
Robert J. Aumann. Agreeing to disagree. The Annals of Statistics , 4(6):1236--1239, 1976
1976
-
[10]
Multi-armed bandit learning in iot networks: Learning helps even in non-stationary settings
R \' e mi Bonnefoi, Lilian Besson, Christophe Moy, Emilie Kaufmann, and Jacques Palicot. Multi-armed bandit learning in iot networks: Learning helps even in non-stationary settings. In Cognitive Radio Oriented Wireless Networks - 12th International Conference, CROWNCOM 2017, L...
2017
-
[11]
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sebastien Bubeck and Nicolo Cesa-Bianchi. Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations and Trends in Machine Learning , 5(1):1--122, 2012
2012
-
[12]
Berry and Bert Fristedt
Donald A. Berry and Bert Fristedt. Bandit problems . Monographs on Statistics and Applied Probability. Chapman & Hall, London, 1985. Sequential allocation of experiments
1985
-
[13]
Bolton and C
P. Bolton and C. Harris. Strategic experimentation. Econometrica , 67(2):349--374, 1999
1999
-
[14]
R. N. Bradt, S. M. Johnson, and S. Karlin. On sequential designs for maximizing the sum of n observations. Annals of Mathematical Statistics , 27:1060--1074, 1956
1956
-
[15]
Distributed multiplayer bandits - a games of thrones approach
Ilai Bistritz and Amir Leshem. Distributed multiplayer bandits - a games of thrones approach. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS) , pages 7222--7232, 2018
2018
-
[16]
Non-stochastic multi-player multi-armed bandits: Optimal rate with collision information, sublinear without
S \' e bastien Bubeck, Yuanzhi Li, Yuval Peres, and Mark Sellke. Non-stochastic multi-player multi-armed bandits: Optimal rate with collision information, sublinear without. CoRR , abs/1904.12233, 2019
1904 arXiv
-
[17]
Matthew Weinberg
Mark Braverman, Jieming Mao, Jon Schneider, and S. Matthew Weinberg. Multi-armed bandit problems with strategic arms. In Conference on Learning Theory, COLT 2019, 25-28 June 2019, Phoenix, AZ, USA , pages 383--416, 2019
2019
-
[18]
Jacobus J. Boomsma. Kin selection versus sexual selection: Why the ends do not meet. Current Biology , 17(16):R673 -- R683, 2007
2007
-
[19]
SIC-MMAB: synchronisation involves communication in multiplayer multi-armed bandits
Etienne Boursier and Vianney Perchet. SIC-MMAB: synchronisation involves communication in multiplayer multi-armed bandits. CoRR , abs/1809.08151, 2018
2018 arXiv
-
[20]
The impact of market structure and learning on the tradeoff between r and d competition and cooperation
David Besanko and Jianjun Wu. The impact of market structure and learning on the tradeoff between r and d competition and cooperation. Journal of Industrial Economics , 61(1):166--201, 2013
2013
-
[21]
Strategic experimentation with exponential bandits
Martin Cripps, Godfrey Keller, and Sven Rady. Strategic experimentation with exponential bandits. Econometrica , 73(1):39--68, 2005
2005
-
[22]
Price of competition and dueling games
Sina Dehghani, Mohammad Taghi Hajiaghayi, Hamid Mahini, and Saeed Seddighin. Price of competition and dueling games. In 43rd International Colloquium on Automata, Languages, and Programming, ICALP 2016, July 11-15, 2016, Rome, Italy , pages 21:1--21:14, 2016
2016
-
[23]
Cooperative and noncooperative research and development in duopoly with spillovers
Claude D'Aspremont and Alexis Jacquemin. Cooperative and noncooperative research and development in duopoly with spillovers. The American Economic Review , 78(5):1133--1137, 1988
1988
-
[24]
Incentivizing exploration
Peter Frazier, David Kempe, Jon Kleinberg, and Robert Kleinberg. Incentivizing exploration. In Proceedings of the Fifteenth ACM Conference on Economics and Computation , EC '14, pages 5--22, New York, NY, USA, 2014. ACM
2014
-
[25]
Subgame-perfect equilibria of finite- and infinite-horizon games
Drew Fudenberg and David Levine. Subgame-perfect equilibria of finite- and infinite-horizon games. Journal of Economic Theory , 31(2):251 -- 268, 1983
1983
-
[26]
Multi-armed Bandit Allocation Indices
John Gittins, Kevin Glazebrook, and Richard Weber. Multi-armed Bandit Allocation Indices . Wiley, 2011
2011
-
[27]
J. C. Gittins. Bandit processes and dynamic allocation indices. Journal of the Royal Statistical Society, Series B , pages 148--177, 1979
1979
-
[28]
Gittins and D.M
J.C. Gittins and D.M. Jones. A dynamic allocation index for the sequential design of experiments. In J. Gani, editor, Progress in Statistics , pages 241--266. North-Holland, Amsterdam, 1974
1974
-
[29]
Christoph Gruter, Ellouise Leadbeater, and Francis L. W. Ratnieks. Social learning: The importance of copying others. Current Biology , 20(16), 2010
2010
-
[30]
W. D. Hamilton. The genetical evolution of social behaviour. i, ii. Journal of Theoretical Biology , 7(1):1--16, 1964
1964
-
[31]
Distributed exploration in multi-armed bandits
Eshcar Hillel, Zohar S Karnin, Tomer Koren, Ronny Lempel, and Oren Somekh. Distributed exploration in multi-armed bandits. In C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 26 , pages 854-...
2013
-
[32]
Strategic experimentation with private payoffs
Paul Heidhues, Sven Rady, and Philipp Strack. Strategic experimentation with private payoffs. Journal of Economic Theory , 159:531--551, 2015
2015
-
[33]
Dueling algorithms
Nicole Immorlica, Adam Tauman Kalai, Brendan Lucier, Ankur Moitra, Andrew Postlewaite, and Moshe Tennenholtz. Dueling algorithms. In Proceedings of the 43rd ACM Symposium on Theory of Computing, STOC 2011, San Jose, CA, USA, 6-8 June 2011 , pages 215--224, 2011
2011
-
[34]
Decentralized learning for multiplayer multiarmed bandits
Dileep Kalathil. Decentralized learning for multiplayer multiarmed bandits. IEEE Transactions on Information Theory , 60(4), 2014
2014
-
[35]
wisdom of the crowd
Ilan Kremer, Yishay Mansour, and Motty Perry. Implementing the "wisdom of the crowd". In Proceedings of the Fourteenth ACM Conference on Electronic Commerce , EC '13, pages 605--606, New York, NY, USA, 2013. ACM
2013
-
[36]
Karlin and Yuval Peres
Anna R. Karlin and Yuval Peres. Game theory, alive . American Mathematical Society, Providence, RI, 2017
2017
-
[37]
Negatively correlated bandits
Nicolas Klein and Sven Rady. Negatively correlated bandits. The Review of Economic Studies , 78(2):693--732, 2011
2011
-
[38]
Vincent Poor
Lifeng Lai, Hai Jiang, and H. Vincent Poor. Medium access in cognitive radio networks: A competitive multi-armed bandit framework. 2008 42nd Asilomar Conference on Signals, Systems and Computers , pages 98--102, 2008
2008
-
[39]
Multiplayer bandits without observing collision information
G \' a bor Lugosi and Abbas Mehrabian. Multiplayer bandits without observing collision information. CoRR , abs/1808.08416, 2018
2018 arXiv
-
[40]
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari. Bandit Algorithms . http://downloads.tor-lattimore.com/banditbook/book.pdf, 2019
2019
-
[41]
Distributed learning in multi-armed bandit with multiple players
Keqin Liu and Qing Zhao. Distributed learning in multi-armed bandit with multiple players. Trans. Sig. Proc. , 58(11):5667--5681, November 2010
2010
-
[42]
Bayesian incentive-compatible bandit exploration
Yishay Mansour, Aleksandrs Slivkins, and Vasilis Syrgkanis. Bayesian incentive-compatible bandit exploration. In Proceedings of the Sixteenth ACM Conference on Economics and Computation , EC '15, pages 565--582, New York, NY, USA, 2015. ACM
2015
-
[43]
Competing bandits: Learning under competition
Yishay Mansour, Aleksandrs Slivkins, and Zhiwei Steven Wu. Competing bandits: Learning under competition. In 9th Innovations in Theoretical Computer Science Conference, ITCS 2018, January 11-14, 2018, Cambridge, MA, USA , pages 48:1--48:27, 2018
2018
-
[44]
Game Theory
Michael Maschler, Eilon Solan, and Shmuel Zamir. Game Theory . Cambridge University Press, 2013
2013
-
[45]
Nowak, Corina E
Martin A. Nowak, Corina E. Tarnita, and Edward O. Wilson. The evolution of eusociality. Nature , 466(7310):1057–1062, 2010
2010
-
[46]
A two-armed bandit theory of market pricing
Michael Rothschild. A two-armed bandit theory of market pricing. Journal of Economic Theory , 9(2):185 -- 202, 1974
1974
-
[47]
Multi-player bandits - a musical chairs approach
Jonathan Rosenski, Ohad Shamir, and Liran Szlak. Multi-player bandits - a musical chairs approach. In Proceedings of the International Conference on Machine Learning (ICML) , pages 155--163, 2016
2016
-
[48]
On games of strategic experimentation
Dinah Rosenberg, Antoine Salomon, and Nicolas Vieille. On games of strategic experimentation. Games and Economic Behavior , 82:31--51
-
[49]
Social learning in one-arm bandit problems
Dinah Rosenberg, Eilon Solan, and Nicolas Vieille. Social learning in one-arm bandit problems. Econometrica , 75(6):1591--1611, 2007
2007
-
[50]
On general minimax theorems
Maurice Sion. On general minimax theorems. Pacific Journal of Mathematics , 8(1):171--176, 1958
1958
-
[51]
Introduction to multi-armed bandits
Aleksandrs Slivkins. Introduction to multi-armed bandits . Foundations and Trends in ML, 2019. draft
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.