{"id":"c473f1b6-c543-47b9-9d8e-00389aa031ee","arxiv_id":"2501.15283","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In a leader/subordinate discussion simulation, most LLM agents do not reproduce the human pronoun pattern, and knowing the pattern does not help them show it.","lead":"Researchers tested whether AI agents playing leaders and subordinates in a team discussion use pronouns the way humans do. They found that most large language models do not reproduce the human pattern, even when the model can state the pattern when asked.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The round-robin long-turn protocol may suppress the conversational contingencies behind the human pronoun asymmetry, so the negative result could be an artifact of simulation design rather than evidence that generative agents cannot reproduce human patterns.","rationale":"The reader's weakest assumption is that the simulation is comparable to the human benchmark, specifically the three-round, round-robin design. I agree that comparability is the central vulnerability, but I sharpen it in two ways. First, the problem is not only the number of rounds but the length and format of each turn: the protocol produces monologues rather than the contingent, reactive exchanges in which the pronoun asymmetry is believed to arise. Second, the paper's inference treats non-significant differences as failures, which converts low statistical power into apparent evidence of absence. Both issues are compounded by the small human effect sizes. The paper does provide some internal support: the explicit-prompt experiment in Appendix B.3 shows that even directly instructing agents to use more \"we\" or \"I\" does not elicit the human pattern, which suggests the failure is not purely a matter of instruction. However, that experiment still uses the same monologue format, so it does not resolve the structural concern. The proposed concrete test would distinguish between 'LLM agents cannot reproduce the pattern' and 'this particular interaction protocol does not elicit the pattern.' Given that the reader's verdict is already CONDITIONAL with a request for sensitivity analyses and a corrected abstract, my analysis reinforces that condition rather than changing the outcome. I do not see grounds to reject the paper or to accept it without the requested robustness checks; the appropriate verdict remains conditional on demonstrating that the result is not an artifact of the simulation protocol.","tokens_in":18914,"tokens_out":8202,"duration_ms":79269,"concrete_test":"Re-run the GPT-4o and Llama-70B conditions under a free-form turn-taking protocol: agents choose when to speak after each utterance, with a controller enforcing one speaker at a time, for the same 41 groups and up to 10 utterances per agent. Apply the same pronoun-frequency and t-test analysis, and additionally report 95% confidence intervals and TOST equivalence bounds anchored at the Kacewicz human effect sizes. If the human-like asymmetry appears under free-form turns, or if CIs exclude the human effect, the conclusion stands; if many gray bars become significant in the human direction, the negative result is an artifact of the monologue protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.4 argues that since pronoun frequency is a ratio, \"the rounds of interactions would not influence our findings.\" This conflates unbiasedness of a proportion with statistical power and ecological validity. With only three round-robin turns per agent, each turn is a long, self-contained monologue (Table 9 shows both leader and non-leader producing polite, collaborative speeches). The human pronoun asymmetries from Kacewicz et al. (2014) emerge in contingent, back-and-forth conversation: lower-status speakers use first-person singular in reactive, tentative talk, while leaders use first-person plural in directive, inclusive talk. A fixed round-robin of monologues removes these speech-act contingencies, potentially suppressing the very signal the paper seeks. Additionally, the evaluation treats any non-significant difference as a failure (gray bars in Figures 2-6). The human effect sizes are small (about 1.3 percentage points for first-person singular, -0.5 for first-person plural), and with 41 groups and roughly 600-2000 words per role, many gray bars may reflect low statistical power rather than the absence of a human-like pattern. The paper's defense in Section 3.4 addresses estimation bias, not variance or content, so the central claim that \"LLM agents barely demonstrate human-like pronoun patterns\" may overreach if the protocol itself prevents the pattern or if null results are underpowered.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether multi-agent LLM simulations reproduce a well-known human result on pronoun usage in hierarchical groups (Kacewicz et al., 2014). The authors simulate 41 four-person groups with one randomly assigned leader and three subordinates, using round-robin, three-turn task-oriented discussions. They test GPT, Llama, Mistral, and Qwen models with four persona prompts from the literature, plus a reflection/planning variant of GPT-4o, and compare the non-leader minus leader frequency differences for first-person singular and first-person plural pronouns against the human benchmark. They find that most model-prompt combinations do not show a statistically significant difference in the human direction; GPT-4o with Prompts 2-4 and GPT-4 with several prompts are exceptions. A separate knowledge probe shows that many of the same LLMs can answer which role tends to use a given pronoun more often, despite failing to reproduce the pattern in interaction. The paper concludes that LLM agents barely resemble human interaction processes and urges caution in LLM-based social simulation.","tokens_in":19143,"tokens_out":9188,"duration_ms":72291,"significance":"If the main result holds, the paper is a useful cautionary contribution to LLM-based social simulation: it is one of the few studies that evaluates an interaction process (pronoun frequencies) rather than only the outcome, it uses an external psychology benchmark from the outset, and it reports full per-condition tables in the appendix, making the evaluation transparent. The knowledge-versus-demonstration gap is a genuine and thought-provoking discrepancy. However, the central negative claim is currently stronger than the evidence, because the simulation protocol may suppress the conversational contingencies that produce the human asymmetry and because null findings are interpreted without a statistical power analysis.","major_comments":[{"comment":"The round-robin protocol in Algorithm 1, with three turns per agent, produces long self-contained monologues rather than the contingent multi-party talk in Kacewicz et al. (2014); Table 9 and Table 10 show that even the non-leader turns are polite collaborative speeches. The argument in §3.4 that 'the frequency does not rely on the number of words generated from each agent' addresses only the unbiasedness of a proportion, not whether the conversational affordances that generate the human asymmetry are present. Since the human effect is theorized to arise in reactive, short back-and-forth exchanges, the negative results on most models may be an artifact of the protocol. I ask the authors to provide evidence that the pronoun asymmetry survives in human data under the same long-turn, fixed-order protocol, or to demonstrate it with a more natural interaction regime before claiming that LLM agents cannot replicate human interactions.","section":"§3.4, Algorithm 1"},{"comment":"The evaluation treats any non-significant difference as a failure (gray bars), but the human effect sizes are small (approximately 1.3 percentage points for first-person singular and approximately -0.5 for first-person plural, Table 4a). With 41 groups and the per-condition variation shown in Appendix B, many of the gray bars may simply reflect low statistical power, and the paper does not report effect sizes, confidence intervals, or a power analysis. The central conclusion that LLM agents barely demonstrate human-like pronoun patterns is therefore stronger than the statistics warrant; a power analysis or an equivalence-testing framework is needed to support the null results.","section":"§4.2, Eq. (1), Figures 2-6"},{"comment":"The abstract and conclusion state that prompt-based and specialized agents 'fail to demonstrate human-like pronoun usage patterns,' but Figure 2a shows GPT-4o with Prompts 2-4 reproducing the human first-person singular pattern, and Figure 5b shows GPT-4 with Prompts 1, 2, and 4 reproducing the human first-person plural pattern. The evidence supports 'most settings fail' but not a categorical failure; the prose should be calibrated to the observed success rate of approximately 8 human-like significant conditions out of 112 model-prompt-pronoun combinations in Table 2.","section":"Abstract, §5.1, §5.3, Conclusion"}],"minor_comments":[{"comment":"Please specify whether the two-sample t-test is paired or unpaired; since leaders and non-leaders come from the same 41 groups, a paired test is the natural choice but is not stated.","section":"§4.2"},{"comment":"The 'May 13th 2025 version' of GPT-4o is inconsistent with the January 2025 submission date and should be corrected.","section":"§4.1"},{"comment":"A sentence explaining how the 'Dem.' counts are derived (number of persona prompts, out of four, with a statistically significant human-like difference) would aid the reader.","section":"§5.4, Table 2"},{"comment":"The Llama 3.1 70B rows list 'Prompt 2' twice, with the second 'Prompt 3' row containing Prompt 2 values; this appears to be a copy-paste error.","section":"Appendix B.2, Table 7"},{"comment":"Calling the forced-choice questionnaire 'know' overstates what is probed; the prompt asks for the aggregate frequency fact, not the context-sensitive sociolinguistic knowledge required for spontaneous use.","section":"§5.4"},{"comment":"The red check marks described in Section 5 are faint in the reproduced figures; please make them more visible or use a different marker.","section":"Figures 2-6"}],"recommendation":"major_revision","confidential_remarks":"The paper is suitable for the journal and the cautionary message is timely. The main risk is that the negative-result claim outruns the simulation design; a validation of the protocol and a statistical reanalysis of the existing data would be sufficient for a revised version. I would also encourage the authors to consider releasing the simulation transcripts and code to support reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a genuinely useful cautionary study for anyone doing LLM social simulation, and it deserves a serious referee, but the abstract oversells the negative result and the simulation protocol is the real soft spot.\n\nThe new thing is concrete: as far as I can tell, this is the first test of whether leader/non-leader pronoun differences from Kacewicz et al. (2014) show up in multi-agent LLM interactions. The coverage is good—several model families, multiple persona prompts, and a reflection/planning agent based on Park et al. (2023). The knowledge-versus-demonstration comparison is the most interesting piece: some models know the pattern when asked directly but do not produce it in the simulated interaction. That is a useful, non-obvious finding. The paper also includes extra experiments (anonymized names, all-female/all-male groups, explicit pronoun instructions) where nothing elicits the human pattern, which strengthens the cautionary message.\n\nThe soft spots are real but not fatal. First, the abstract says agents 'fail' when GPT-4o with prompts 2-4 actually shows the human first-person-singular pattern. That is an overclaim; the accurate headline is 'most configurations do not replicate the pattern.' Second, the simulation uses three rounds of round-robin monologue turns. Human pronoun asymmetries come out most clearly in contingent, short-turn conversation, and the paper's claim that round count doesn't matter because frequency is a ratio conflates unbiased estimation with ecological validity. The negative result may partly be an artifact of the protocol. I'd want to see a sensitivity analysis with more turns or free-form discussion before accepting the strong version of the conclusion. Third, no code or raw data is released, so the numbers cannot be independently recomputed. That's a standard ask these days.\n\nStatistical power is a minor additional concern—the human effect sizes are small—but the fact that some significant results do appear (e.g., GPT-4o) suggests the method has some sensitivity, so I wouldn't call this a fatal flaw.\n\nBottom line: the paper is for practitioners and researchers evaluating LLM-based social simulation. It deserves peer review, but only with revisions: fix the abstract, release data/code, and address the protocol's ecological validity. I'd be happy to engage with it.","headline":"A useful cautionary study on LLM social simulation, but the abstract overclaims and the round-robin protocol may suppress the signal worth studying.","tokens_in":19701,"tokens_out":2633,"would_cite":true,"duration_ms":24315,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM agents barely reproduce human pronoun patterns in simulated team interactions, even when they know the patterns.","keywords":["LLM agents","social simulation","pronoun usage","hierarchical interaction","persona prompts","multi-agent interaction","generative agents","leader and non-leader roles"],"falsifier":"A replication that lets four LLM agents hold a free-form, 30-minute discussion of the same ranking task and finds a statistically significant positive Δ for first-person singular pronouns (non-leaders above leaders) and negative Δ for first-person plural across most models and prompts would contradict the paper's central claim; so would a controlled test showing that the three-round round-robin format itself suppresses human-like pronoun patterns when humans speak under the same constraint.","tokens_in":18608,"feed_emoji":"💬","tokens_out":4943,"duration_ms":39586,"temperature":0.7,"pith_summary":"The paper asks whether interactions among LLM agents resemble human interactions, using a well-documented marker: in task-oriented group discussions, non-leaders use first-person singular pronouns (\"I\", \"me\") more often than leaders, while leaders use first-person plural pronouns (\"we\", \"us\") more often. Replicating the four-person, ranking-task setup of a known psychology study with 41 simulated groups, the authors find that across multiple model families, persona prompts, and a reflection/planning agent, LLM agents barely demonstrate these human-like pronoun patterns. Even when an LLM can correctly answer a direct question about who uses which pronoun more often, it fails to enact that pattern in its generated dialogue. If correct, this means practitioners cannot treat LLM-based social simulations as trustworthy stand-ins for human group processes.","feed_headline":"LLM agents barely reproduce human pronoun patterns in teams","feed_subtitle":"Models that know leaders say \"we\" more still don't show it in simulated group talk, so social simulations can mislead.","key_machinery":"The load-bearing device is the leader/non-leader pronoun-frequency differential, Δ = f_nonleaders,avg − f_leaders,avg, computed for first-person singular and first-person plural pronouns across 41 simulated four-person groups; paired with four persona prompts from prior work and a reflection/planning agent adapted from Park et al. (2023), it converts a well-known human interaction norm into a quantitative, falsifiable benchmark for agent behavior.","core_discovery":"The paper's central claim is that LLM agents barely demonstrate human-like pronoun patterns, even if the LLM agent may show some understanding of those patterns. Replicating the four-person, task-oriented ranking discussion of Kacewicz et al. (2014) with 41 simulated groups, the authors measure the difference in pronoun frequency between non-leaders and leaders, Δ = f_nonleaders,avg − f_leaders,avg, for first-person singular and first-person plural pronouns. Contrary to human results, most non-leader LLMs do not use first-person singular pronouns more often, and leader LLMs do not use first-person plural pronouns more often; in many cases the trend is reversed, and almost none of the tested models, prompts, or the reflection/planning specialized agent yield statistically significant human-like differences. The paper also establishes a knowledge–demonstration gap: several LLMs, when asked directly and with permuted answer orders, can identify the correct human pattern but fail to exhibit it in their interaction process.","pith_inferences":["If the gap generalizes beyond pronouns, other unconscious conversational signals such as turn-taking, interruption, politeness, and hedging may also fail to emerge in LLM agent interactions, so interaction-level validation should become standard before simulation results are used.","The three-round, round-robin design may compress the dynamics that let human leaders and subordinates negotiate roles; longer or free-form conversations, or letting agents choose turn order, might allow status roles to crystallize and should be tested before concluding the failure is intrinsic.","The knowledge–demonstration gap suggests a testable intervention: prompting agents to explicitly monitor their own pronoun use or reflect on their role after each turn might close the gap, though the paper's planning/reflection results hint the opposite, pointing to a dissociation between declarative social knowledge and procedural language generation.","Another testable extension is to measure whether the same failure appears consistently across random seeds for a fixed model and prompt, which would separate genuine role-behavior deficits from generation variance."],"forward_implications":["Conclusions drawn from LLM social simulations about group interaction processes may contradict established human findings, so practitioners should not rely on these simulations for decisions about real social dynamics.","Adding reflection and planning components does not close the gap; for GPT-4o it actually made pronoun patterns less human-like than the simple prompt-based agent.","Model family and prompt choice, not model size or capability, dominate pronoun patterns: Llama and Qwen families show internally consistent but non-human trends, and larger models are not more human-like.","The knowledge–demonstration gap means that an LLM's ability to state a social norm is not evidence that its multi-agent interactions embody that norm.","Even explicit instructions to use certain pronouns more often (e.g., \"Please use first-person plural forms more often\" for leaders) did not elicit the human pattern."],"supporting_citations":[{"why":"Supplies the human pronoun-usage benchmark and the four-person task-oriented group setup that the simulation replicates.","marker":"Kacewicz et al. (2014)"},{"why":"Provides the generative-agent architecture with memory, reflection, and planning that the specialized agent adapts.","marker":"Park et al. (2023)"},{"why":"Documents the GPT model family (GPT-3.5, GPT-4, GPT-4o) used as agents in the experiments.","marker":"Achiam et al. (2023)"},{"why":"Provides the Llama 3.1 family of models used in the experiments.","marker":"Dubey et al. (2024)"},{"why":"Provides the Qwen 2.5 family of models used in the experiments.","marker":"Bai et al. (2023)"},{"why":"Provides the Mistral model used in the experiments.","marker":"Jiang et al. (2023a)"},{"why":"Motivates the permutation of answer orders in the knowledge-probing prompt, making the knowledge–demonstration comparison robust.","marker":"Zheng et al. (2023)"}],"fun_headline_variants":["LLM agents fail to mimic human pronoun use in team talks","Even LLMs that know the pattern don't show it in simulations","Simulated groups show LLMs can't copy human pronoun habits","AI agents miss human pronoun cues in group discussions","LLM social simulations fall short on pronoun usage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulation is a fair stand-in for the human study: three rounds of round-robin turns among four agents produce pronoun frequencies comparable to the original 30-minute human group discussions.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents fail to mimic human pronoun use in team talks","Even LLMs that know the pattern don't show it in simulations","Simulated groups show LLMs can't copy human pronoun habits","AI agents miss human pronoun cues in group discussions","LLM social simulations fall short on pronoun usage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00059,"raw_usage":{"total_tokens":2740,"prompt_tokens":891,"completion_tokens":1849,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":1768}},"tokens_in":507,"tokens_out":1849,"duration_ms":10023,"temperature":1.0,"reasoning_tokens":1768,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:26:09.890443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that lets four LLM agents hold a free-form, 30-minute discussion of the same ranking task and finds a statistically significant positive Δ for first-person singular pronouns (non-leaders above leaders) and negative Δ for first-person plural across most models and prompts would contradict the paper's central claim; so would a controlled test showing that the three-round round-robin format itself suppresses human-like pronoun patterns when humans speak under the same constraint.","supporting_citations":[],"review_version":1}