{"id":"e4a66e99-c9b7-458b-82ef-ec3361e462e8","arxiv_id":"2412.03359","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An open platform and benchmark for evaluating LLM-based multi-agent systems through the 'Who is Spy?' game, including a leaderboard and behavioral analysis of ten models.","lead":"This paper presents WiS, an online platform where LLM-powered agents play the social deduction game 'Who is Spy?' to test their reasoning, deception, and resistance to manipulation. It also publishes a leaderboard and experimental results suggesting that GPT-4o is a strong reasoner while Qwen models excel at deception.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"WiS's core claim that its game metrics measure reasoning, deception, and adversarial robustness is unvalidated: no external benchmark, human baseline, or confound control (word-pair difficulty, game count) supports the construct, so Table 2's GPT-4o 'superiority' may be a game artifact rather than…","rationale":"I read the paper as a systems/infrastructure contribution with an empirical demonstration. The platform itself is real, publicly accessible, and includes a data-download API, SDK, visualization, and a large live competition (15,893 rounds, 200 participants, Appendix A)—these are genuine independent supports for usability and scalability. The paper also makes falsifiable predictions (e.g., Table 2 rankings, Table 3 attack/defense effects) that could be checked against logs. However, the central claim in the abstract and conclusion—that WiS 'effectively distinguishes' attacking, defense, reasoning, and deception capabilities—requires that the game metrics validly measure those constructs. That condition is the least secure. The scoring rules in Section 3 are a reasonable game design, but the mapping from score to 'reasoning' or 'deception' is stipulated, not demonstrated. There is no external benchmark or human baseline; the reasoning experiment changes prompts and interprets outcome differences as differences in internal reasoning quality; and the leaderboard scoring explicitly rewards play volume. Appendix A's model-substitution findings further muddy real-world leaderboard validity. I therefore agree with the reader's weakest-assumption identification. The right verdict remains CONDITIONAL: accept the infrastructure contribution, but require construct-validity evidence (external calibration, confound-controlled re-analysis, or human baseline) before the capability-ranking claims are taken as established. My concern does not move the verdict; it sharpens the condition.","tokens_in":11194,"tokens_out":8121,"duration_ms":80730,"concrete_test":"Run a calibration cohort: select 10–15 additional LLMs with published scores on established reasoning and theory-of-mind benchmarks (MMLU, GPQA, ARC, ToMi), run each through the Section 6.1 WiS protocol, and compute Spearman rank correlations between WiS average score / civilian vote accuracy and each external benchmark. If correlations are weak or confidence intervals include 0, the claimed construct validity fails; if strong, it is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that WiS 'effectively distinguishes' attacking, defense, reasoning, and deception capabilities—depends on treating the Section 3 scoring rules and Section 5.1 indicators (average score, vote accuracy, foul rate) as valid measures of those constructs. The paper provides no external anchor: no correlation with established reasoning or theory-of-mind benchmarks, no human baseline, and no ablation that isolates one capability from another. Section 5.2's definitions are asserted rather than validated ('voting accuracy is the most relevant metric for assessing analytical reasoning'), and Section 6.2 then explains GPT-4o's higher average score as 'attributed to its enhanced reasoning abilities'—a circular inference. The scoring system itself conflates ability with luck: spy/civilian word-pair difficulty and speaker order are uncontrolled, and the leaderboard's cumulative score explicitly rewards number of games played ('the more games played, the more likely one is to achieve a high ranking'). The paper's own Appendix A.3 further shows that real leaderboard entries often use different models than declared, so the leaderboard cannot cleanly rank model capabilities. Section 6.4's reasoning experiment is also self-referential: adding a 'reasoning' prompt changes other agents' behavior, and the authors attribute GPT-4o's improvement and Qwen/Llama's decline to unmeasured 'reasoning quality.' These weaknesses do not disprove the platform's usefulness, but they mean the headline capability-ranking claim is not supported by the evidence as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces WiS, an online platform that uses the \"Who is Spy?\" social deduction game to evaluate LLM-based multi-agent systems. It provides a Hugging Face/API integration interface, a live leaderboard, game visualization, data download, and custom scoring rules. The authors run games among ten LLM agents and report overall performance (Table 2), prompt-injection attack/defense experiments (Table 3), and a reasoning-prompt experiment (Figure 2), concluding that the platform effectively distinguishes attacking, defense, reasoning, and deception capabilities and that GPT-4o is superior. The platform and code are publicly released.","tokens_in":11523,"tokens_out":7133,"duration_ms":63197,"significance":"If its evaluation constructs are valid, the platform would be a useful, scalable, and less overfit-prone complement to static benchmarks, with the additional benefit of releasing game logs for training. The engineering contribution is concrete: public deployment, an SDK, visualization, and downloadable data are real strengths. However, the paper's central empirical claim is not yet supported because the metrics are not externally anchored, sample sizes are small, and several controlled-variable and identity issues affect the leaderboard and reasoning experiments. The contribution is therefore more convincingly a platform/tool description than a validated evaluation benchmark at this stage.","major_comments":[{"comment":"The central claim that WiS 'effectively distinguishes' attacking, defense, reasoning, and deception capabilities is not yet supported because the metric construct is unvalidated. Section 5.1 asserts that average score reflects 'comprehensive abilities' and that voting accuracy is 'the most relevant metric for assessing an agent's analytical reasoning ability,' but no correlation with an established reasoning or theory-of-mind benchmark, no human baseline, and no capability-isolating ablation is provided. In addition, word-pair difficulty and starting-speaker order are uncontrolled, so GPT-4o's advantage in Table 2 and the capability attributions in Section 6.2 could be artifacts of game setup rather than measurements of the intended constructs. I ask for an external validation step (e.g., correlation with existing benchmarks or human expert ratings) or controlled ablations that isolate each claimed capability.","section":"Section 5.1 and Table 2"},{"comment":"The statistical evidence is too weak for the cross-model distinctions claimed. Appendix C says each role-condition was repeated 'more than 24 times,' but the Table 3 values include win rates such as 33.33%, 18.75%, and 23.53%, which are consistent with win counts as small as 8 out of 24, 3 out of 16, and 4 out of 17; even if the real denominators are larger, the paper never states exact counts. No confidence intervals or significance tests are reported anywhere, and because each six-agent game couples all participants' outcomes, the effective number of independent observations is smaller still. Please report exact game counts, per-condition confidence intervals, and appropriate statistical tests, and show that the sample size supports the specific rankings claimed in Section 6.2.","section":"Appendix C and Table 3"},{"comment":"The evaluated model set is described inconsistently. Appendix C states that 'ten publicly available open-source models' were evaluated, but Table 2 includes GPT-4o, Gemini-1.5-pro, Claude-3-5-Sonnet, Kimi, ERNIE, and Doubao, which are closed-source or API-only models. This matters for reproducibility because API-backed models change over time and because the paper's scope claims differ between 'open- and closed-source LLMs' in the abstract and 'open-source models' in the appendix. Please specify exact model versions, access mode (weights versus API), and evaluation dates, and correct the inconsistent characterization.","section":"Section 6.1 and Appendix C"},{"comment":"The leaderboard is not a clean model-capability ranking. The scoring rule in Section 3 states that 'the more games played, the more likely one is to achieve a high ranking,' and Appendix A.3 reports that top entries sometimes used models different from the declared ones (e.g., a GPT-3.5 entry actually using Doubao). The leaderboard therefore conflates engagement and identity verification with model quality. Please restrict leaderboard rankings to verified model identities with game-count-controlled scores, or present the leaderboard explicitly as a community-competition ranking rather than as evidence of model capability.","section":"Section 3 and Appendix A.3"},{"comment":"The reasoning experiment is confounded by its own intervention. Adding the 'Reasoning' prompt from Section 5.2 changes the content and style of public speech by the target civilian, and other agents observe that speech; the measured changes in voting accuracy and win rate can therefore reflect information-propaganda effects rather than the reasoning quality of the prompted model. Attributing GPT-4o's improvement to 'superior chain-of-thought reasoning' and Qwen's and Llama's decline to 'relatively weak reasoning' (Section 6.4) is circular because reasoning quality is not measured independently. Please redesign the experiment so that reasoning is elicited privately (e.g., before the public statement) or add a control condition that matches utterance length and style.","section":"Section 6.4 and Figure 2"}],"minor_comments":[{"comment":"The sum 'NX i=1 si' should be typeset as \\sum_{i=1}^N s_i, and N should be defined explicitly as the number of games played by that player.","section":"Section 3, displayed equation"},{"comment":"Model names are inconsistent (e.g., 'GPT4o' vs. 'GPT-4o' and 'Qwen2.5-72B-Instruct' vs. 'QWEN'), and the game name appears both as 'Who is Spy?' and 'Who is spy'; please standardize these terms.","section":"Throughout"},{"comment":"The reference list contains duplicate entries for Xu et al. 2023 with the same title ('Exploring large language models for communication games: An empirical study on werewolf'), and the in-text citations to 'Xu et al.' do not distinguish the three different Xu et al. works; please disambiguate.","section":"Section 2 and References"},{"comment":"The example phrases 'Bitter taste. From tree. Keep awake.' are not explained and look like sample descriptions rather than word-pair examples; please label them or remove them.","section":"Figure 1"},{"comment":"The paper does not specify how the attacking, defense, and reasoning prompts are inserted (system message, user message, or appended context), nor how they interact with the base system prompt; this detail is needed for reproducibility.","section":"Section 5.2, Table 1"},{"comment":"The statements 'each experiment was repeated over 90 times' and 'more than 24 times' should be reconciled and made precise about whether these are independent games per model per condition.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The platform is a real, public infrastructure contribution and the paper ships code and data, which I value. However, the empirical claims currently overreach the evidence, and the evaluation section needs substantial revision rather than cosmetic changes. I would be comfortable with acceptance only after the construct-validity and statistical-rigor issues are addressed; the paper may also be better framed as a systems/demo contribution until then."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper ships a real system, and that is the main thing worth knowing. The WiS platform is live, has a public leaderboard, downloadable game logs, and the Who-is-Spy environment is a lightweight social deduction game that is cheaper to run than Werewolf or Avalon. If you want to compare LLM agents in a low-cost interactive setting, this is a plausible download-and-go option. The attack/defense prompt setups and the data release are also genuinely new pieces of infrastructure.\n\nThe authors cite the existing social deduction literature properly and do not claim to invent the genre. The platform design is straightforward: unified model API, leaderboard, visualization, data download. That is fine for a systems contribution.\n\nThe soft spots are in the empirical validation, not the artifact. Table 2 reports win rates and scores with no error bars or significance tests; across 90 games per model with six-player dynamics, those numbers are noisy. Table 3's attack/defense results include win rates like 33.33% that likely come from a handful of games. More importantly, the construct validity is unvalidated: the paper asserts that voting accuracy measures reasoning, average spy score measures deception, and foul rate measures attack effectiveness, but there is no external benchmark, no human baseline, and no ablation that isolates one capability. Section 6.4's reasoning experiment is also confounded because adding the reasoning prompt changes what other agents hear, so GPT-4o's improvement could be coordination or influence, not reasoning ability.\n\nThe public leaderboard has its own issues: the cumulative score explicitly rewards playing more games, and Appendix A.3 shows real users substitute models under the hood. The authors report that substitution rather than hiding it, which is to their credit, but it means the live leaderboard cannot cleanly rank model capabilities.\n\nThese problems do not sink the platform as a tool. You can still use it to gather data and observe behavior. But the title's \"Enhancing Evaluation\" is only as strong as the metrics, and the metrics need validation before anyone should trust the capability rankings. If the authors add a small human baseline or correlate with established reasoning/ToM benchmarks, or simply reframe the paper as a platform contribution without the strong capability-ranking claims, it becomes a solid systems paper.\n\nWho is this for? People building or testing LLM agents who want a free, interactive environment with real opponents and logged play, and the agent-evaluation community as a data source. Not a must-read for theory-minded folks.\n\nI would send this to a serious referee, but with the instruction to push hard on metrics validation. The artifact is real and the authors are honest about some of its warts, so referee time is warranted.","headline":"A live, useful platform for game-based LLM-agent evaluation, but the capability-ranking claims run ahead of the evidence.","tokens_in":12073,"tokens_out":2160,"would_cite":false,"duration_ms":21951,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 'Who is Spy' game can rank LLM agents by reasoning, deception, and attack resistance.","keywords":["LLM multi-agent systems","Who is Spy","social deduction games","game-based evaluation","deception detection","reasoning evaluation","prompt injection attack","leaderboard platform"],"falsifier":"Run the same ten models on a standard set of logic and reasoning puzzles and compare the ordering with their WiS average scores; if the orderings are uncorrelated, or if a human panel judging game transcripts cannot distinguish the agents the platform ranks, the claimed capability separation fails.","tokens_in":11019,"feed_emoji":"🕵️","tokens_out":8042,"duration_ms":73435,"temperature":0.7,"pith_summary":"This paper introduces WiS, an online platform that evaluates large language model agents by having six of them play the social deduction game 'Who is Spy'. The paper's central claim is that performance in this game, measured by a custom zero-sum scoring system, voting accuracy, win rates, and foul rates, can distinguish agents' attacking, defense, reasoning, and deception abilities. To support this, the authors ran repeated games among ten open- and closed-source models and report clear behavioral differences, with GPT-4o achieving the highest overall win rate and average score while Qwen models showed the strongest spy win rate, suggesting stronger deception. The platform is open, scalable, and built for continuous leaderboard updates, with game data downloadable for further model training. If the claim holds, the game provides a dynamic, less easily gamed alternative to static datasets for evaluating multi-agent systems.","feed_headline":"GPT-4o tops 'Who is Spy' benchmark for LLM agents","feed_subtitle":"A live arena where six AI agents play spy versus civilians, scoring reasoning, deception, and attack resistance.","key_machinery":"The load-bearing object is the 'Who is Spy' game environment together with its scoring rules and role-specific metrics. The game fixes six participants, a spy and five civilians, each with a word they must describe without revealing; rounds of description, fouling, and voting decide elimination, and the spy wins by surviving to round three or leaving fewer than three civilians. The scoring system is designed as a zero-sum allocation of 12 points per game, giving the spy 0, 4, 8, or 12 points depending on when discovery happens, plus one bonus point per round in which civilians vote out the spy, so total payoff is fixed. These rules turn win rate, average score, voting accuracy, survival rounds, and foul rate into the quantitative evidence from which the paper reads reasoning, deception, and attack/defense capabilities.","core_discovery":"On its own terms, the paper establishes that a six-player 'Who is Spy' game can serve as a multi-agent evaluation environment that ranks LLM-based agents by distinct capabilities. Each round, one spy and five civilians describe related but different words without saying them, vote to eliminate a player, and score points under a zero-sum rule: 12 points go either to a surviving spy or to the surviving civilians when the spy is found, with smaller payoffs for later discovery and bonus points for correct votes. The paper argues that spy-role scores measure deception, civilian voting accuracy measures reasoning, and responses to inserted prompt-injection instructions measure attacking and defense. In tests repeated over ninety games per model, GPT-4o had the highest civilian win rate (84.93%), overall win rate (76.67%), and average score (3.24), while Qwen2.5-72B-Instruct had the highest spy win rate (46.60%). The authors interpret these and related results as evidence that the platform effectively differentiates multi-agent abilities.","pith_inferences":["If the game score is meant to measure general capabilities, it should be validated against established reasoning and deception benchmarks; the paper does not provide that external check, so transferability remains an open question.","The zero-sum scoring rewards survival timing and vote precision more than raw win frequency, so average score and win rate can rank agents differently; treating them as one number would blur the two.","The appendix shows that top competitors sometimes used different models than declared and added defensive filters, which implies the live leaderboard partly measures prompt engineering and hardening, not just base-model ability.","Using the game directly as a training objective could encourage overfitting to spy-game speech patterns; whether those skills transfer to other adversarial or cooperative settings is untested."],"forward_implications":["The platform yields a continuously updated, open leaderboard that ranks LLM agents by their game scores without needing new static datasets.","The prompt-injection attack and defense settings give a quantitative comparison of how easily different models can be manipulated into fouls or bad votes.","Adding an explicit reasoning step changes performance unequally: it raises GPT-4o's voting accuracy and civilian win rate while lowering those of Qwen2.5-72B-Instruct and Llama-3-70B-Instruct.","Downloadable game logs support supervised or reinforcement learning, so the same environment can be used for evaluation and for training improved agents.","Because any model can be registered through the unified interface, new agents can be compared in real time against current best performers."],"supporting_citations":[{"why":"Supplies the stated motivation: LLM multi-agent environments suffer from instability and reproducibility problems that WiS targets.","marker":"(Guo et al., 2024)"},{"why":"Example of the game/outcome-based multi-agent evaluation line that the WiS environment extends with finer-grained metrics.","marker":"(Hong et al., 2023)"},{"why":"Representative debate-environment evaluator that the paper contrasts with game-based evaluation.","marker":"(Chan et al., 2024)"},{"why":"GameEval, the conversational-game benchmark that provides the closest existing comparison for game-based LLM evaluation.","marker":"(Qiao et al., 2023)"},{"why":"Avalonbench, a social deduction benchmark the paper distinguishes from its multi-dimensional behavioral analysis.","marker":"(Light et al., 2023)"},{"why":"Empirical Werewolf study showing social deduction games can reveal LLM communication and deception skills.","marker":"(Xu et al., 2023a)"},{"why":"Avalon 'game of thoughts' work on deception that motivates using deduction games to measure deceptive behavior.","marker":"(Wang et al., 2023b)"}],"fun_headline_variants":["GPT-4o wins 'Who is Spy' agent arena","Spy game ranks LLM agents: GPT-4o on top","'Who is Spy' arena: GPT-4o leads LLM agents","New game benchmark pits LLM agents: GPT-4o wins","LLM agents duel in 'Who is Spy' – GPT-4o best"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument's load-bearing premise is that the custom 'Who is Spy' score is a valid proxy for general reasoning, deception, and adversarial robustness; the paper offers no external benchmark, human baseline, or independent reasoning test to establish that mapping.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o wins 'Who is Spy' agent arena","Spy game ranks LLM agents: GPT-4o on top","'Who is Spy' arena: GPT-4o leads LLM agents","New game benchmark pits LLM agents: GPT-4o wins","LLM agents duel in 'Who is Spy' – GPT-4o best"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1315,"prompt_tokens":966,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":254}},"tokens_in":582,"tokens_out":349,"duration_ms":3567,"temperature":1.0,"reasoning_tokens":254,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:29:03.144737+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same ten models on a standard set of logic and reasoning puzzles and compare the ordering with their WiS average scores; if the orderings are uncorrelated, or if a human panel judging game transcripts cannot distinguish the agents the platform ranks, the claimed capability separation fails.","supporting_citations":[{"cited_title":"Chawla, Olaf Wiest, and Xiangliang Zhang","cited_arxiv_id":null,"evidence_quote":"Supplies the stated motivation: LLM multi-agent environments suffer from instability and reproducibility problems that WiS targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Avalonbench, a social deduction benchmark the paper distinguishes from its multi-dimensional behavioral analysis."}],"review_version":1}