{"id":"de15bce7-73ea-4fd6-a4ce-f65cfbd630c0","arxiv_id":"2608.09128","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A new 21-game benchmark scores LLM social reasoning with rule-decided outcomes, and SPaRTan, a self-reflection loop, transfers playbook lessons across games.","lead":"This paper introduces Social Gym, a suite of 21 multiplayer games where AI agents are scored by objective game outcomes instead of subjective judges. It also tests SPaRTan, a training-free loop in which an agent writes itself a strategy guide from its own matches and uses it in later games.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SPaRTan's 'lifts the weaker side' claim is under-identified: n=30 cells are mostly within the paper's own ±18pp CI, and without a placebo-playbook control the observed shifts cannot be attributed to reflection content.","rationale":"I read the paper in good faith: the Social Gym environment is a substantial engineering contribution, the rule-decided outcomes genuinely remove LLM-judge subjectivity, the code release is a real reproducibility asset, and the self-play decomposition in Appendix I is a thoughtful way to separate role balance from capability gaps. The reader's conditional verdict is appropriate. My concern targets the second half of the central claim, the SPaRTan improvement. The paper itself identifies the two weakest points in its Limitations: small per-condition samples and the absence of a placebo-playbook control. That makes this a missing-evidence problem rather than a speculative one. Several headline deltas are individually within the paper's own confidence intervals, and the no-placebo design cannot separate content effects from generic prompt-perturbation effects. Since the paper concludes that a single regularity emerges across all four setups, this is load-bearing: if a placebo reproduces the side-asymmetric pattern, the method's contribution reduces to 'inserting text changes behavior,' not to self-play-derived playbooks. I do not think this warrants rejection, because the benchmark contribution stands and the SPaRTan effect may survive a proper test; it warrants the conditional verdict the reader already gave, with the placebo and larger-sample controls as explicit acceptance conditions. My agreement with the reader is partial: the reader's stated weakest assumption was the Appendix H game simplifications, whereas I see the causal identification of the SPaRTan effect as the more decisive risk, though the reader's rationale does flag the same small-sample and placebo issues.","tokens_in":31844,"tokens_out":12931,"duration_ms":138883,"concrete_test":"Run one decisive experiment on the core cells of §5.1 and §5.2 (Werewolves R1/R3 alt, Resistance Rwcsu alt, Spyfall R1/R2 alt, plus their main-side counterparts) with n=100 per condition, in four arms: vanilla baseline, SPaRTan playbook, a length-matched placebo (scrambled playbook or unrelated strategy prose not derived from trajectories), and a no-injection prompt-length control. If the alt-side gain over vanilla does not exceed the placebo gain by at least about 10pp, or if the main-side drop appears in the placebo arm, then the content-specific 'lift weaker side' claim is not supported. Pre-register the primary comparison before running.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim has two halves: the verifiable benchmark and the training-free improvement. The benchmark half is reasonably supported by the released engine, rule-decided outcomes, and the self-play/cross-play decomposition in Appendix I. The improvement half is not yet supported. To establish that SPaRTan 'lifts the structurally weaker side,' the win-rate shifts must be (i) real rather than sampling noise and (ii) caused by the playbook's content rather than by any inserted prompt text. The paper's own Limitations section admits that per-condition n=30 gives binomial 95% CIs of roughly ±18pp, and that there is no placebo-playbook control. In Table 1, the averaged alt-side gain from baseline to R3 is 24%→43%, a 19pp shift that is not significant at the 95% level for two n=30 proportions (z≈1.6, p≈0.11). The cross-game transfer in Figure 3 has median +7pp, far inside noise. Individual peaks such as Werewolves alt 23%→63% are larger, but the headline 'single regularity' is assembled from many underpowered cells without multiple-comparison control, and Figure 4 contains a salient counterexample: Undercover alt regresses from 17% to 0%. The paper explicitly says in Limitations that several effects are within the noise range and should be replicated, yet the conclusion still states that a single regularity emerges. That overstates the evidence. A length-matched placebo arm is needed to rule out a generic prompt-perturbation reading, which is especially important because the observed side-asymmetric pattern could in principle arise from any instruction that changes the behavior of one side of an asymmetric game.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Social Gym, an environment of 21 multi-agent social games with rule-decided outcomes, and uses a Bradley–Terry/Elo tournament to benchmark seven LLMs. The leaderboard roughly tracks general model capability but shows large per-game and per-role inversions. The paper then proposes SPaRTan, a training-free play–reflect–transfer loop in which a model writes a first-person strategic playbook from its own self-play trajectories and injects it into the system prompt for later games. Across within-game iteration, cross-game transfer, multigame transfer, and distillation experiments, the authors report that GPT-5-mini playbooks tend to lift the structurally weaker side of asymmetric hidden-role games, while Qwen3-32B shows mostly null results except in action-channel games.","tokens_in":32186,"tokens_out":4233,"duration_ms":44241,"significance":"If the benchmark half is accepted, Social Gym is a useful, verifiable complement to LLM-judge-based social evaluation: it provides rule-decided outcomes, a unified Elo procedure with bootstrap confidence intervals, released code, and a thoughtful decomposition of role balance versus capability using self-play. The SPaRTan claim would also be valuable as a training-free improvement method. However, the improvement half is currently under-supported: the core experiments use n=30 per condition, most reported deltas fall within the paper's own stated ±18 pp binomial confidence interval, and there is no placebo-playbook control. The benchmark contribution is solid and reproducible; the SPaRTan contribution needs substantially stronger evidence or a more hedged interpretation.","major_comments":[{"comment":"The central improvement claim is not supported at the stated confidence level. Each condition has n=30 games, giving binomial 95% CIs of roughly ±18 pp, as the Limitations section admits. The averaged alt-side shift from 24% (baseline) to 43% (R3) is a 19 pp change, which is not significant at the 95% level for two independent n=30 proportions (z≈1.6, p≈0.11); individual cells in Table 1 are almost all within the noise band, and Figure 4 contains a salient counterexample (Undercover alt falling from 17% to 0%). The paper reports no multiple-comparison control across the many cells. Larger sample sizes, per-cell confidence intervals, or a pre-registered aggregate test are needed before 'a single regularity emerges' can be stated.","section":"§5.1, Table 1; Limitations"},{"comment":"The attribution of the observed win-rate shifts to the playbook's content is not identified. The design compares R-armed players against vanilla opponents, but not against opponents given a length-matched, content-matched placebo (e.g., scrambled or unrelated instructions). Without such a control, the effects in Table 1 and Figure 3 could be generic prompt-perturbation effects or changes in verbosity or assertiveness rather than the specific strategic content of the reflection. The paper itself acknowledges this missing control in the Limitations; a placebo arm in at least the two flagship conditions (Werewolves and Resistance) is necessary to support the claimed mechanism.","section":"§5.1 and Limitations (No placebo-playbook control)"},{"comment":"The construct validity of the benchmark requires additional evidence because of the listed modifications: Skull ends after one round (wins_needed=1), Sheriff of Nottingham is reduced to a binary honest/smuggle choice with fixed payoffs, Werewolves is capped at 40 turns, and Coup block claims resolve automatically. The paper asserts that these games still exercise deception, negotiation, and coalition tracking, but provides no evidence that the simplified dynamics retain those demands rather than collapsing to pattern-matched heuristics. At minimum, the authors should report ablation or manipulation checks (e.g., whether expert-level playbooks or known strategies produce the expected performance shifts, or how often games terminate at the 40-turn cap) to show that the modified implementations measure the intended social-cognitive skills.","section":"Appendix H (Game Suite Details)"},{"comment":"The cross-game and multigame evidence is too underpowered to support the sign-flip interpretation. In Figure 3, after excluding the saturated Chameleon column, the alt-side cells have median +7 pp and the main-side cells median -7 pp, both well inside the ±18 pp per-cell noise band; the multigame results in Figure 4 show three of four non-saturated targets either not beating Single-R1 or degrading, with the only clear positive cell being held-out Resistance at +24 pp. The conclusion that playbooks transfer across games and 'lift the weaker side' therefore rests on sign patterns over noisy cells rather than on statistically separable effects. Either aggregate the cells with a mixed-effects model or report adjusted tests.","section":"§5.1 (Cross-game transfer) and §5.1 (Multigame transfer)"}],"minor_comments":[{"comment":"The table header reads 'BL/R0' but the caption initially says 'BL/R 0'; define R0 explicitly in the caption (the text defines it in §5.1 only later).","section":"Table 1"},{"comment":"The left panel's category separators are easy to miss; labeling each category block or adding a second header row would help readers map columns to the five game categories.","section":"Figure 2"},{"comment":"The PD row reports three values (alt/main/tie); the main text describes a 13% to 58% lift, but it should state that PD's baseline includes 74% mutual-cooperate ties and that win rate is strict wins only, to avoid confusion.","section":"§5.3.1, Table 10"},{"comment":"The phrase '40K context window' should be '40k context window' for consistency.","section":"Appendix M.4"},{"comment":"Some references have typographical artifacts (e.g., 'V oyager'), and the table in Appendix K labels 'Gemini' without disambiguating it as 'Gemini 3.1 Pro' in the caption; a final proofread is needed.","section":"References and Appendix K"},{"comment":"The method name is typeset inconsistently as 'SPARTAN' in the title and 'SPARTan' in the abstract; pick one spelling.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The benchmark contribution is solid and worth publishing: the environment, rule-decided outcomes, Elo methodology, and released code are reproducible, and the self-play/cross-play decomposition in Appendix I is a genuine analytical strength. The SPaRTan claim, however, is the centerpiece of the paper and is currently under-powered and lacks a placebo control; the authors' own Limitations section concedes both points. I would not reject, because the requested revisions are feasible within the manuscript's scope, but the conclusion should either be substantially hedged or supported by additional experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the benchmark half is real and worth taking seriously; the improvement half is not yet established. This is a paper with a reproducible, rule-scored 21-game environment, a clean Elo pipeline, and a genuinely underpowered SPaRTan analysis whose headline regularity is overstated.\n\nWhat is new: Social Gym fills a real gap. SOTOPIA-style judge-based evaluation has known biases, and single-game Werewolf or Avalon studies do not generalize. Twenty-one games in one engine with rule-decided outcomes and released code is a genuine step forward. The self-play/tournament decomposition in Appendix I is thoughtful: cross-play role gaps conflate capability and role balance, and same-model self-play separates them. The qualitative parroting evidence in Appendix J is credible and adds texture. The paper is also unusually honest in its Limitations section, which is a real strength.\n\nThe soft spots are where the stress-test lands. SPaRTan's central claim rests on small samples. The paper itself says n=30 gives binomial 95% CIs of roughly ±18 pp and that several effects are within that noise, yet the conclusion still says a single regularity emerges. The headline averaged alt-side gain in Table 1, 24% to 43%, is not significant at the 95% level for two n=30 proportions (z about 1.6). The cross-game transfer median of +7 pp is inside noise. Figure 4 contains Undercover alt regressing from 17% to 0% under R. There is no placebo-playbook control, so the observed shifts cannot be attributed to reflection content rather than generic prompt perturbation. The authors acknowledge this too. Per-game Elos also lack error bars. None of this sinks the benchmark, but the regularity should be presented as a hypothesis, not a result.\n\nI do not think the game simplifications in Appendix H are the core problem. They are disclosed, and the suite is broad enough that the environment still exercises the target skills. The external-validity worry is real but secondary to the statistical one.\n\nThe Qwen3-32B null results are a useful contrast: Prisoner's Dilemma works, most discussion-heavy games do not. That action-channel versus free-discussion pattern is plausible and worth reporting, but it too is post-hoc and underpowered.\n\nWho this is for: anyone building multi-agent social evaluation or studying training-free self-improvement. The benchmark deserves a serious referee and likely acceptance after revision. The SPaRTan sections need larger n, per-game confidence intervals, a placebo control, and a toned-down conclusion. I would cite the benchmark, not the method, in the next year.","headline":"A reproducible 21-game benchmark that deserves uptake; the SPaRTan 'single regularity' overstates what n=30 cells and no placebo control can support.","tokens_in":32753,"tokens_out":2931,"would_cite":true,"duration_ms":60489,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Social Gym and SPaRTan show LLM social reasoning can be measured and improved without LLM judges.","keywords":["Social Gym","SPaRTan","multi-agent social games","LLM social reasoning","verifiable evaluation","self-play reflection","Elo tournament","hidden-role deduction"],"falsifier":"Replay the 21 games under full unmodified commercial rules; if the leaderboard inversions and SPaRTan weak-side lifts disappear, the simplifications in Appendix H are responsible for the results.","tokens_in":31631,"feed_emoji":"🎲","tokens_out":7162,"duration_ms":68598,"temperature":0.7,"pith_summary":"This paper argues that LLM social reasoning can be measured objectively by staging 21 multi-agent games whose winners are decided by game rules, not by an LLM judge, and ranked in an Elo tournament. On this benchmark no model is uniformly strong: per-game rankings invert sharply even though overall Elo tracks general capability. The paper further claims that a training-free loop, SPaRTan, in which a model plays self-play games, writes a first-person strategic playbook from its own trajectories, and injects that playbook into its system prompt, lifts the structurally weaker side of asymmetric hidden-role games for a strong closed model, transfers across games, and distills to weaker models. The same loop largely fails for an open-weights model on discussion-heavy games, succeeding only when the required action collapses to a single token. If these claims hold, social reasoning can be benchmarked reproducibly and improved without weight updates.","feed_headline":"21-game arena measures LLM social skills without judges","feed_subtitle":"Per-game rankings invert; a self-reflection playbook lifts weaker sides in asymmetric games.","key_machinery":"The load-bearing mechanism is the pairing of a rule-decided game engine with the SPaRTan play–reflect–transfer loop. The engine is a finite-state machine with partial observability: messages are filtered as public, team-private, or private, so agents must reason from incomplete information, and every episode ends in an algorithmic win/loss/score that feeds a Bradley–Terry Elo fit. SPaRTan takes a model's own self-play trajectories, asks it to write a first-person, game-agnostic strategic playbook covering deception, detection, persuasion, information management, coalition dynamics, and timing, and prepends that playbook to the system prompt of future games. The playbook is the transferable object: it carries reflection-derived strategy across games and into weaker models without any weight update.","core_discovery":"The central discovery is that a rule-decided, multi-game tournament exposes role- and game-specific weaknesses that overall leaderboards hide, and that a self-generated natural-language playbook can partially close those weaknesses. In Social Gym, 21 games spanning normal-form, economic, bluffing, hidden-role deduction, and social strategy are run on a shared engine, and outcomes are aggregated into per-game and overall Bradley–Terry Elo ratings. GPT-5-mini tops the overall leaderboard, but per-game Elos invert, and in same-model self-play the minority/deceptive side of Werewolves, Spyfall, and Resistance wins only 23%, 20%, and 30% of games for the strongest model. SPaRTan's iterated reflections raise that alt-side win rate to peaks of 63%, 46%, and 40% respectively while lowering the majority side, and the same playbook transfers across games and to weaker students on the disadvantaged side. The effect is capacity-dependent: Qwen3-32B shows no gains on free-discussion games, only on Prisoner's Dilemma and partially Resistance, whose decisive actions are single discrete tokens.","pith_inferences":["The lift-the-weaker-side regularity suggests a general principle: reflection helps a model most where its own baseline is structurally disadvantaged; a testable extension would apply SPaRTan to asymmetric human-agent tasks such as negotiation or customer-service de-escalation, where role asymmetry is explicit.","The contrast between Qwen3-32B's clean gains on Prisoner's Dilemma and flat results on free-discussion games implies that the bottleneck is the model's ability to execute strategy through a long dialogue channel; constraining the action space or adding structured reasoning may make self-reflection effective for smaller models.","Without a content-matched placebo playbook, part of the observed gains could be generic prompt-perturbation effects; a scrambled-playbook control would separate content from instruction-following.","Several benchmark games have saturated baselines, such as Chameleon alt-side at 100% for strong models, so the current suite may need difficulty calibration to discriminate among frontier models."],"forward_implications":["Overall Elo rankings should not be read as a single social-intelligence score; per-game and per-role Elo expose inversions that a scalar hides.","For strong models, a training-free reflection loop can partially close the gap on the structurally weaker side of asymmetric games, within-game and across held-out games.","Playbooks generated by a strong model can be injected into weaker models and improve them on the disadvantaged side, so strategies learned in games can be shared without fine-tuning.","Because every Social Gym episode yields a rule-computed reward, the environment is directly usable as a training signal for reinforcement learning with verifiable rewards, replacing LLM judges.","The null result on an open-weights model implies that self-reflection gains are capacity- and action-channel-dependent, not a universal property of LLMs."],"supporting_citations":[{"why":"Supplies the base interaction loop that Social Gym generalizes to N-agent games.","marker":"Zhou et al., 2024b"},{"why":"Provides the repeated matrix-game setup and results that Social Gym's normal-form games extend.","marker":"Akata et al., 2025"},{"why":"Shows that Werewolf can yield verifiable rule-decided outcomes, the key idea Social Gym generalizes.","marker":"Xu et al., 2023"},{"why":"Supplies the Bradley–Terry Elo aggregation methodology used for the leaderboard.","marker":"Chiang et al., 2024"},{"why":"Supplies the paired-comparison statistical model behind the Elo fits.","marker":"Bradley and Terry, 1952"},{"why":"Defines the verbal self-reflection approach that SPaRTan adapts into a transferable playbook.","marker":"Shinn et al., 2023"},{"why":"Supplies the skill-library line that SPaRTan's cross-game playbook transfer extends.","marker":"Zhao et al., 2024"}],"fun_headline_variants":["Social Gym: 21 games, no judges, Elo rankings for LLMs","Rule-decided games expose LLM social weaknesses; playbook patches","Self-play reflection playbook improves LLM social game performance","Tournament reveals uneven LLM social skills; self-play helps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's validity rests on the assumption that the simplified game implementations in Appendix H, such as Skull ending after one round, Sheriff reduced to a binary smuggle choice, the 40-turn Werewolves cap, and automated Coup blocks, still exercise the deception, negotiation, and coalition-tracking skills the paper claims to measure rather than short-circuiting them.","fun_headline_variants_meta":{"raw":{"variants":["Social Gym: 21 games, no judges, Elo rankings for LLMs","Rule-decided games expose LLM social weaknesses; playbook patches","Self-play reflection playbook improves LLM social game performance","Tournament reveals uneven LLM social skills; self-play helps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001078,"raw_usage":{"total_tokens":4562,"prompt_tokens":1045,"completion_tokens":3517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":3451}},"tokens_in":661,"tokens_out":3517,"duration_ms":25127,"temperature":1.0,"reasoning_tokens":3451,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:58:00.380687+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replay the 21 games under full unmodified commercial rules; if the leaderboard inversions and SPaRTan weak-side lifts disappear, the simplifications in Appendix H are responsible for the results.","supporting_citations":[],"review_version":1}