{"id":"6f917581-d8bd-4b80-adb2-b23a41a9e410","arxiv_id":"1908.01135","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a two-player bandit game where players see each other's actions but not rewards, competition reduces exploration, cooperation increases it, and neutral players can outperform a single player by observing each other.","lead":"This paper studies how competition or cooperation between two players changes how much they explore a risky option in a bandit learning problem. It shows competitors explore less than a single player, cooperators explore more, and neutral players can learn from each other.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4 Part 2 proof has two unproved steps: Bob's zero-sum subgame advantage is only shown for p>(mβ+g)/(1+β), and Alice's payoff under Bob's deviation is assumed ≥ single-player optimum without justification.","rationale":"The reader's weakest assumption correctly identifies the subgame gap: Theorem 1's proof only supports Bob's winning copycat strategy for p > (mβ+g)/(1+β), whereas the theorem statement only gives p* ≤ (mβ+g)/(1+β). That gap is real and directly affects Theorem 4 Part 2. I also find a second, independent gap in the same proof: the lower bound E(Γ'_A) ≥ α/(1−β) for Alice's payoff when Bob unilaterally deviates to S'_B is asserted without argument and is not a consequence of perfect Bayesian equilibrium, since Alice's strategy is only optimal against S_B. This makes the proof of the central 'learn from each other' claim incomplete even if the subgame property holds. These are proof gaps rather than demonstrated falsehoods, so the appropriate verdict remains conditional acceptance pending repair. The paper's main ideas are plausible and the other theorems (thresholds for competition/cooperation, long-run convergence) are supported by more substantial arguments, including concentration-based reasoning in Section 6. I credit the paper for the explicit thresholds and the unifying framework, but the main novelty claim in Theorem 4 Part 2 needs a complete proof.","tokens_in":32237,"tokens_out":15815,"duration_ms":161220,"concrete_test":"Re-derive Theorem 4 Part 2 using only inequality (8) from Theorem 1's proof, without citing Theorem 1 itself. For a uniform prior, β→1, and p=0.65 (in the gap (2−√2, (mβ+g)/(1+β)]), compute the value of the zero-sum subgame with Alice forced to R and Bob to L in round 0 and optimal continuation; if Bob's maxmin net advantage is not positive for all p in the gap, the subgame claim is false. Separately, construct a candidate PBE in the neutral game where Alice's strategy free-rides on Bob's informative play, and evaluate whether E(Γ'_A) under Bob's deviation can fall below α/(1−β); if it can, the asserted lower bound is invalid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 4 Part 2 (Section 5) is incomplete. Bob's deviation S'_B requires that in the zero-sum subgame where Alice is forced to play R at round k and Bob plays L, Bob can guarantee a strictly positive net advantage for every p in (p*,g). The text cites 'By Theorem 1', but Theorem 1 only proves p* ≤ (mβ+g)/(1+β); its copycat strategy gives Bob a strict advantage only when p > (mβ+g)/(1+β), leaving p in (p*, (mβ+g)/(1+β)] unsupported. Separately, the proof asserts 'Since the equilibrium is perfect Bayesian, we have E(Γ'_A) ≥ α/(1−β)' for the profile (S_A,S'_B). This does not follow: S_A is a best response to S_B, not to the deviating S'_B, and Bob's uninformative left-until-exploration play can destroy information S_A exploited. Without this lower bound, E(Γ'_B)>E(Γ'_A) does not imply Bob's deviation beats α/(1−β). Both steps are load-bearing for the claim that neutral players strictly outperform the single-player optimum.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a two-player one-armed bandit in which, each round, a player chooses a predictable left arm (known success probability p) or a risky right arm (prior μ), observes own reward and the other player's action but not the other player's reward, and has utility Γ_i + λΓ_j. The three regimes are competing (λ=-1), neutral (λ=0), and cooperating (λ=1). The main claims are: competing players explore less than a single player for p above a threshold p* ≤ (mβ+g)/(1+β), yet still explore for p below about m+βw/2 (Theorems 1 and 2); cooperating players explore for some p above the Gittins index g (Theorem 3); neutral players explore with probability one for p<g and, in every perfect Bayesian equilibrium for p∈(p*,g), each player strictly beats the single-player optimum (Theorem 4); and competing and neutral players eventually settle on the same arm in every Nash equilibrium while cooperating players may oscillate (Theorems 9-10 and Proposition 2). Finite-horizon analogues and improved bounds for a uniform prior are also given.","tokens_in":32460,"tokens_out":17657,"duration_ms":167784,"significance":"If the proofs are completed, this is a valuable contribution to multiplayer bandit learning and strategic experimentation. The λ-interpolation gives a clean comparative framework, and the paper makes concrete falsifiable predictions: competition reduces exploration, cooperation increases it, neutral players can profit from observing each other's actions, and long-run agreement fails only outside the [-1,0] range. The paper also has genuine technical strengths: the copycat strategy is explicit, the concentration lemma (Lemma 4) is clean, and the thresholds are defined intrinsically rather than fitted to data. The main caveat is that the proof of the headline welfare result for neutral players, Theorem 4 Part 2, currently rests on two unjustified steps, and the algebra in Theorem 1 needs correction. These are fixable in principle, but they are load-bearing.","major_comments":[{"comment":"The invocation 'By Theorem 1' does not cover all p∈(p*,g). Theorem 1's proof gives Bob a strict advantage via the copycat strategy only when p > (mβ+g)/(1+β), and the theorem only establishes p* ≤ (mβ+g)/(1+β). Thus the interval p∈(p*, (mβ+g)/(1+β)] is left unsupported. A separate argument is needed showing that in the zero-sum subgame where Alice is forced to play R at round k, Bob can guarantee a strictly positive net advantage for every p in (p*,g).","section":"Section 5, proof of Theorem 4 Part 2"},{"comment":"The assertion 'Since the equilibrium is perfect Bayesian, we have E(Γ'_A) ≥ α/(1−β)' does not follow. Alice's strategy S_A is a best response to S_B, not to the deviating strategy S'_B; under S'_B, which plays left until Alice explores, Alice may lose information that she exploited in the original equilibrium, so her payoff could fall below the single-player optimum α/(1−β). Without this lower bound, the inequality E(Γ'_B)>E(Γ'_A) does not imply that Bob's deviation beats α/(1−β), which is the contradiction the proof needs.","section":"Section 5, proof of Theorem 4 Part 2"},{"comment":"The algebraic identity leading to inequality (8) is incorrect. From the displayed expressions for E(Γ_A) and E(Γ_B) one obtains E(Γ_A)-E(Γ_B) = (m-p(1+β))β^k + (1-β)∑_{t=k+1}^∞ E(γ_A(t))β^t, not the displayed expression with an extra factor β^k on the tail sum. Consequently inequality (8) does not follow as stated. A corrected derivation changes the no-exploration threshold to (m+βg)/(1+β) (under the same bound on the tail), and Theorem 3's use of (8) inherits the problem. The stated theorems may still be true, but the proof must be redone.","section":"Section 3, Eqs. (7)-(8)"}],"minor_comments":[{"comment":"The model restricts λ to [-1,1], but Proposition 3 analyzes λ<-1; please clarify whether that proposition is intended as an out-of-model remark or whether the model should allow λ outside [-1,1].","section":"Section 1.1 and Section 6.1"},{"comment":"The definition of p* as sup{p: arm R is explored in some Nash equilibrium} already makes 'for all p>p* the players do not explore' true by definition; the content of the theorem is the upper bound p*≤(mβ+g)/(1+β). The statement could be rephrased to avoid this redundancy.","section":"Section 1.2.1, Theorem 1"},{"comment":"The displayed chain from the equilibrium bound to the inequality Φ_k ≥ (α−p)/(1−p−β^k) is compressed; the term handling the case where exploration has already occurred is omitted in the first displayed inequality and only appears implicitly in the next line. Please spell out the derivation.","section":"Section 5, proof of Theorem 4 Part 1"},{"comment":"The sentence 'there is a perfect Bayesian equilibrium in which Bob visits both arms infinitely often whenever' ends abruptly; the trailing 'whenever' should be removed or completed.","section":"Section 6.1, Proposition 3"},{"comment":"Several small typos remain: 'Rotschild' should be 'Rothschild' in Section 7, and 'Salomon' in the bibliography entry [RSV] should be 'Solan'. The reference [RSV] also lacks a year.","section":"References and typos"}],"recommendation":"major_revision","confidential_remarks":"The paper has a strong conceptual core and several correct-looking auxiliary results, but the two unproved steps in Theorem 4 Part 2 are exactly the steps needed for the paper's most striking welfare claim, and the algebra in Theorem 1 needs correction. I would encourage the authors to supply the missing subgame argument and the missing PBE payoff bound; with those in place the paper would be a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The λ-interpolating setup is the thing to remember: one model, with λ = -1, 0, 1, and clean qualitative answers—competition suppresses exploration, cooperation amplifies it, neutral players can learn from each other. The paper earns credit for that framing. The copycat strategy in Theorem 1 is elegant, and the long-term convergence proofs in Section 6 are substantial, especially Lemma 4 and the good/bad-history machinery. Theorems 2 and 3 and the finite-horizon analogues are also genuinely useful. The neutral-convergence result overlaps with Aoyagi and HRS15, as the authors acknowledge, but the strict-improvement result and the zero-sum analysis go beyond those papers. This is a serious theory paper, not a repackaging.\n\nThe soft spots are in Theorem 4 Part 2, and they are load-bearing. The proof says 'By Theorem 1' to claim Bob wins the zero-sum subgame after Alice's first exploration for all p in (p*, g). But Theorem 1's construction only gives Bob a strict advantage for p > (mβ + g)/(1 + β), a number that can be strictly larger than p*. The interval between is unsupported. Separately, the proof asserts E(Γ'_A) ≥ α/(1 − β) under (S_A, S'_B) because the equilibrium is perfect Bayesian. That does not follow: S_A is a best response to S_B, not to Bob's deviation; Bob's uninformative play can destroy information that Alice's equilibrium strategy exploited. Without that lower bound, E(Γ'_B) > E(Γ'_A) does not imply Bob's deviation beats the single-player benchmark. Both steps are needed for the headline claim that neutral players strictly outperform a single player. This needs fixing, or the claim needs to be restricted.\n\nMinor but real: the statements of Theorems 1 and 3 assert 'with probability 1' no exploration, but the proofs compare expectations. There is probably a zero-sum optimality argument that bridges the gap—if an action yields negative expected payoff against a particular opponent response, it cannot be used with positive probability in an optimal strategy—but the text never makes that explicit.\n\nOverall: take the framework and the qualitative results as good news. Do not quote Theorem 4 Part 2 until the subgame property and the payoff bound are proved. The paper deserves a serious referee; I would send it out, with instructions to focus on Section 5. If the authors close those two gaps, this is a solid, citable contribution to the strategic experimentation literature.","headline":"A useful unifying framework with mostly solid qualitative results, but Theorem 4 Part 2 has two real proof gaps and the almost-sure claims outrun the argument; worth refereeing, not yet citable as stated.","tokens_in":32990,"tokens_out":3094,"would_cite":true,"duration_ms":33014,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that in a two-player one-armed bandit, competition lowers the exploration threshold below the solo Gittins index, cooperation raises it above, and neutral players can each beat the solo optimum by observing each other's…","keywords":["multi-armed bandits","strategic experimentation","exploration-exploitation tradeoff","zero-sum games","cooperative games","Gittins index","Nash equilibrium","perfect Bayesian equilibrium"],"falsifier":"Compute, for a fixed prior $\\mu$ and discount $\\beta$, the value of the zero-sum game that starts with Alice forced to play the risky arm in round 0 and Bob forced to play the known arm. If for some $p\\in(p^*,g)$ Bob's value is not strictly positive, the pivot of the neutral-learning theorem fails. A simulation counterpart is to truncate the game at a large horizon, solve for the perfect Bayesian equilibrium by dynamic programming, and check whether each player's expected reward still exceeds the single-player optimum on that interval.","tokens_in":32020,"feed_emoji":"🎰","tokens_out":10230,"duration_ms":103378,"temperature":0.7,"pith_summary":"This paper studies a two-player, one-armed bandit game in which each player sees the other's arm choices but not the other's rewards. It asks how the exploration-exploitation tradeoff changes when the players are competing, cooperating, or neutral, and it answers with three threshold statements relative to the Gittins index of the risky arm. Competing players stop exploring above a threshold below the solo optimum; cooperating players explore above the solo optimum; and neutral players, in a middle range of arm quality, each earn strictly more than a single optimal player because they can infer information from each other's actions. The paper also proves that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while cooperating players need not.","feed_headline":"Bandit rivals explore less, partners more, neutrals win","feed_subtitle":"A two-player bandit analysis shows when rivalry freezes exploration, when teamwork encourages it, and when watching wins.","key_machinery":"The load-bearing objects are the Gittins index $g=g(\\mu,\\beta)$, the threshold known-arm probability at which a single player is indifferent between the known and risky arms; the copycat strategy, in which a player stays on the known arm until the opponent explores and then repeats the opponent's previous move one round later; and the induced threshold $p^*\\leq (m\\beta+g)/(1+\\beta)$ where copying destroys the value of exploration. The cooperative result uses a lagged-copy arrangement that gives the team two observations per experiment. The neutral learning result is carried by a perfect Bayesian deviation argument: if a player received only the single-player optimum, she could switch, upon the opponent's first exploration, to the zero-sum copycat strategy and win in the remaining subgame, contradicting equilibrium. The long-term convergence results use a concentration inequality on the empirical mean of the risky arm to show that an infinitely exploring player eventually identifies the better arm and that the other player follows.","core_discovery":"The paper's central claim is that the value of information in the two-player game is alignment-dependent. In the zero-sum regime ($\\lambda=-1$) there is a threshold $p^*<g$, with the copycat bound $p^*\\leq (m\\beta+g)/(1+\\beta)$, such that for every $p>p^*$ neither player ever explores in equilibrium; a player who experiments first hands the opponent a one-round-lagged copy of the information, and above $p^*$ that erases the explorer's advantage. Below the threshold $m+\\beta w/2$, however, both players explore in the first round, so competition does not reduce play to pure myopia. In the fully cooperative regime ($\\lambda=1$), one player can explore while the other copies with a delay, effectively turning one experiment into two observations, and this makes the team explore for some $p>g$. In the neutral regime ($\\lambda=0$), for $p$ between the competing threshold $p^*$ and the solo Gittins index $g$, every perfect Bayesian equilibrium gives each player strictly more expected reward than a single player using an optimal strategy; the mechanism is that after the opponent's first exploration, a player can switch to a winning zero-sum strategy against her. Finally, in every Nash equilibrium competing and neutral players converge to the same arm with probability 1, whereas cooperating players have equilibria with infinitely many switches.","pith_inferences":["If the copycat bound can be sharpened to cover the whole interval $(p^*,g)$, the neutral-learning theorem would pin the end of mutually beneficial learning exactly at the competition threshold; a direct numerical check is to compute the value of the zero-sum subgame after a forced exploration.","The same copycat mechanism suggests that any rule raising the cost of imitation, such as a temporary protection for the arm a player first explored, would widen the exploration region for competing and neutral players.","The exponential decay of the no-exploration probability gives a quantitative handle for algorithm designers: a finite-time agent can estimate equilibrium exploration probabilities and decide when observation has made further own-experimentation unnecessary.","The non-convergence example for cooperating players indicates that aligned payoffs alone do not ensure coordinated specialization; equilibrium selection or communication would be needed to make teams settle on a single arm."],"forward_implications":["Competing players will not explore the risky arm for any $p$ above $p^*$, so head-to-head rivalry can freeze experimentation even when a solo learner would continue.","Competing players still explore for all sufficiently small $p$, so the zero-sum interaction does not collapse to always playing the safer arm.","Two cooperating players can explore for values of $p$ where a single player would stop, because a lagged-copy arrangement makes one experiment yield two observations.","Neutral players who observe actions but not rewards can each beat the single-player optimum in every perfect Bayesian equilibrium for $p\\in(p^*,g)$, and the probability that no one explores decays exponentially.","In every Nash equilibrium, competing and neutral players settle on the same arm almost surely, while cooperating players can have equilibria with one player switching arms infinitely often."],"supporting_citations":[{"why":"Defines the Gittins index and the single-player exploration threshold that all three regimes are compared against.","marker":"[BJK56]"},{"why":"Shows the index's role in multi-armed bandit allocation, grounding the single-player benchmark $g$.","marker":"[GJ74]"},{"why":"Supplies the minimax theorem used to give the zero-sum game a value and justify optimal competing strategies.","marker":"[Sio58]"},{"why":"Raises the question of whether players settle on the same arm, which the paper's convergence theorems answer for competing and neutral players.","marker":"[Rot74]"},{"why":"Solves convergence under the same imperfect monitoring in related bandit models, providing the comparison point for the paper's convergence theorems.","marker":"[Aoy98]"},{"why":"Guarantees existence of perfect Bayesian equilibria, which the neutral learning theorem requires.","marker":"[FL83]"}],"fun_headline_variants":["Rivalry curbs bandit exploration, partnership fuels it","In bandit duels, competition stalls, cooperation sparks exploration","Neutral bandits observe, learn, and outplay soloists","Competing bandits halt, cooperating bandits explore more","Bandit rivals stand still, partners explore, neutrals triumph"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The neutral-learning theorem rests on the unproved subgame property that after one player explores, the other can guarantee a strictly positive expected advantage for every $p$ in $(p^*,g)$; the paper's copycat argument proves this only for $p$ above the larger cutoff $(m\\beta+g)/(1+\\beta)$, so the interval between the two cutoffs is where the claim is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Rivalry curbs bandit exploration, partnership fuels it","In bandit duels, competition stalls, cooperation sparks exploration","Neutral bandits observe, learn, and outplay soloists","Competing bandits halt, cooperating bandits explore more","Bandit rivals stand still, partners explore, neutrals triumph"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000721,"raw_usage":{"total_tokens":3379,"prompt_tokens":1230,"completion_tokens":2149,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":846,"completion_tokens_details":{"reasoning_tokens":2063}},"tokens_in":846,"tokens_out":2149,"duration_ms":17889,"temperature":1.0,"reasoning_tokens":2063,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:28:18.029637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, for a fixed prior $\\mu$ and discount $\\beta$, the value of the zero-sum game that starts with Alice forced to play the risky arm in round 0 and Bob forced to play the known arm. If for some $p\\in(p^*,g)$ Bob's value is not strictly positive, the pivot of the neutral-learning theorem fails. A simulation counterpart is to truncate the game at a large horizon, solve for the perfect Bayesian equilibrium by dynamic programming, and check whether each player's expected reward still exceeds the single-player optimum on that interval.","supporting_citations":[],"review_version":1}