REVIEW 4 major objections 13 references
Personality shifts reshape how multi-agent LLM teams talk everywhere, but only harm objective outcomes when the task lacks a structured artifact that can absorb the damage.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Low agreeableness massively shifts multi-agent LLM communication yet barely hurts coding milestones, while the same prompt sharply degrades research milestones and collapses bargaining agreements.
T0 review reviewed 2026-07-15 challenge →
load-bearing objection Solid multi-domain empirical result: same low-A shift leaves coding milestones mostly intact but tanks research and bargaining; the artifact-buffer story is useful but only partly isolated. the 4 major comments →
When Does Personality Composition Matter for Multi-Agent LLM Teams?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The authors establish that personality effects on multi-agent LLM performance are task-contingent. The same low-agreeableness manipulation produces large communication-state shifts across models and domains, yet milestone outcomes stay near baseline in structured coding for three of four models while research milestones drop substantially and bargaining acceptance falls to near zero. High agreeableness is largely null. They attribute the pattern to artifact-mediated buffering: formal code constraints filter process degradation; unconstrained text and concession decisions do not.
What carries the argument
Artifact-mediated buffering, tracked by communication state φ (the share of acts that explore via questions, disagreements, and suggestions rather than converge via acknowledgments). Structured code deliverables constrain the solution space independently of discourse quality; unstructured research ideas and accept/offer decisions do not, so the same φ shift reaches outcomes only in the latter domains.
Load-bearing premise
That domain differences in outcome sensitivity are caused mainly by whether the deliverable is a formal constrained artifact, rather than by other things that also differ across the domains, such as agent count, interaction rules, evaluation method, or leftover effects of loaded wording.
What would settle it
Equalize confounds and re-run the same low-agreeableness and neutral prompts on matched coding versus research protocols, or on a competitive high-structure task: if coding milestones then degrade like research, or if the buffer fails once agent count and judging are matched, the artifact-structure claim fails.
If this is right
- Multi-agent system designers should treat personality composition as task-structure-dependent, not as a universal team setting.
- Negatively valenced personality adjectives inflate adversarial profiles; neutral behavioral descriptors yield smaller, more model-specific effects.
- Low agreeableness is better used as a bounded lead-position critic than as a team-wide trait.
- High-agreeableness prompting adds little under current training priors; personality prompting mainly induces adversarial behavior.
- Process metrics (planning quality, hostility, φ) can diverge from final milestones and code quality in structured domains.
Where Pith is reading between the lines
- Production coding fleets may tolerate disagreeable sub-agents for milestone success, while research and negotiation fleets cannot.
- The untested competitive-plus-high-structure cell could still break or preserve buffering and is the natural next stress test.
- Residual research degradation under neutral wording for some models implies safety-training differences still matter after valence is controlled.
- Artifact-only evaluation will systematically under-detect toxic multi-agent dynamics whenever the final deliverable is formally constrained.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether personality prompting (primarily Big Five agreeableness via Goldberg bipolar adjectives) changes objective multi-agent LLM team outcomes, or only communication style. Across frontier models and three MultiAgentBench domains—cooperative high-structure coding, cooperative low-structure research, and competitive bargaining—low agreeableness reliably shifts communication state φ toward exploration (often ~0.44→0.93), with trait ablations showing the shift is largely agreeableness-specific. Outcome effects are task-contingent: coding milestones and LLM-judged code quality stay near baseline for most models despite degraded planning quality, while research milestones drop substantially and bargaining agreement collapses (often to ≤1%). A neutral-paraphrase control attenuates but does not fully erase effects; a single-challenger pilot suggests role/position-dependent heterogeneous composition. The authors interpret the pattern as artifact-mediated buffering: formal code constraints insulate outcomes from process degradation, whereas unstructured outputs and concession decisions do not.
Significance. The work addresses a timely design question for production multi-agent systems (Claude Code-style orchestration, AutoGen/MetaGPT-class frameworks): when personality composition is a useful control knob versus noise. Strengths include multi-model coverage, three-domain comparison, process–outcome dissociation with an explicit communication measure (φ, δ), conscientiousness/openness ablations, second-judge agreement on act labels, bootstrap CIs, bargaining offer-movement vs never-accept controls, and a neutral-paraphrase experiment that partially disentangles trait from prompt valence. If the task-contingent pattern holds under cleaner isolation, the paper supplies a concrete, falsifiable design principle—structured deliverables buffer adversarial process noise; open-ended and concession-based tasks do not—and a methodological caution about negatively loaded personality adjectives. That combination is of clear interest to multi-agent systems and LLM evaluation venues.
major comments (4)
- §5.1, Fig. 1, Table 2: The central causal attribution—that coding robustness is due to artifact structure buffering process degradation—is under-isolated. Coding (3 agents, structured code actions, 5 tasks), research (5 agents, free-form ideas, 15 tasks), and bargaining (2 agents, structured actions, 50 products) co-vary in agent count, protocol, deliverable type, and evaluation surface. Table 5 and §4.3/§5.3 already show that under neutral wording Grok-3 and DeepSeek research remain degraded while GPT-4o largely recovers, so residual valence and model-specific factors still explain part of the dissociation. Either (i) add a within-domain structure manipulation (e.g., same team protocol with vs without format-enforced parsing / formal acceptance criteria), or (ii) reframe the main claim as a descriptive task-contingent pattern and demote artifact-mediated buffering to a hypothesis, with
- §4.2, Table 2, Table 9, Appendix H: The coding null claim (milestones largely unchanged for 3/4 models) is load-bearing for the buffering story but rests on a thin task set (5 collaborative game-development tasks) and mixed results (Grok-3 d=1.69, p=0.017; GPT-4o/DeepSeek directionally down but n.s.). Bootstrap CIs for Claude/GPT-4o/DeepSeek milestones include zero but are wide. Please report power or equivalence tests for the null, increase coding task diversity/N, and avoid language that treats coding as uniformly buffered when one frontier model shows a clear milestone drop.
- §3.2 and §4.2: Milestone evaluation uses the same LLM-as-judge rubric for coding and research, which is good for cross-domain comparison, but research milestones are themselves unconstrained natural-language quality judgments and may be more sensitive to adversarial tone than code milestones that can be partially grounded in executable structure. Without human validation or an alternative automatic metric for research ideas, part of the research degradation could be judge sensitivity to low-A discourse rather than true idea quality. A small human-rated subsample or dual-judge protocol for research milestones (analogous to Appendix B for acts) is needed to support the claim that outcomes—not only process scores—degrade.
- §5.4, Appendix C, Table 7: Design recommendations about lead-position challengers are drawn from a small single-challenger pilot with incomplete cells and no formal interaction tests (position × domain × model). Present this strictly as exploratory, remove or heavily hedge prescriptive language in the conclusion, and do not let heterogeneous composition carry weight equal to the homogeneous main results until a fuller factorial is available.
Circularity Check
Empirical multi-domain measurement study; outcomes are external task metrics, not forced by the communication definition or by fitted inputs.
full rationale
This paper does not present a first-principles derivation whose conclusions reduce to its inputs by construction. Personality prompts (Goldberg bipolar adjectives at fixed qualifier levels) are experimental treatments, not parameters fitted to the outcome claims. The communication state φ is an operational ratio of labeled acts (questions, disagreements, suggestions over those plus acknowledgments); milestones and bargaining agreement rates are separate MultiAgentBench task metrics. Reporting that low-A raises φ while coding milestones stay near baseline for 3/4 models is a measured dissociation, not a tautology: φ does not enter the milestone definition, and code quality scores are independently LLM-judged. Trait ablations (C, O), neutral-paraphrase controls, second-judge act labels, and never-accept bargaining logs are additional measurements, not self-citation chains that force the task-structure claim. Citations (MultiAgentBench, Serapio-García et al., Bell, Peeters et al.) supply protocols and framing; none is a uniqueness theorem or ansatz by the present authors that makes the cross-domain pattern true by definition. Interpretive language about artifact-mediated buffering is post-hoc explanation of empirical patterns and is open to confounds (agent count, protocol, residual valence), but that is a causal-isolation concern, not circularity. Score 0: no circular step of the enumerated kinds.
Axiom & Free-Parameter Ledger
free parameters (3)
- Goldberg qualifier intensity levels (primary: 2 vs 8 on a 9-level scale)
- Task set sizes (5 coding tasks; 15 research tasks; 50 bargaining products)
- Communication classifier model and temperature (GPT-4o-mini, T=0)
axioms (4)
- domain assumption Goldberg bipolar adjective markers, combined with linguistic qualifiers, validly shape Big Five-like traits in frontier LLMs for multi-agent interaction.
- domain assumption MultiAgentBench milestone counts and bargaining agreement rates are comparable objective success measures across domains under a shared LLM-as-judge rubric for coding/research.
- ad hoc to paper Bales-inspired multi-label acts (question, disagreement, suggestion, acknowledgment) and the exploration fraction φ adequately capture process degradation relevant to team outcomes.
- ad hoc to paper Structured code deliverables constrain the solution space independently of communication quality sufficiently to buffer milestone and code-quality outcomes.
invented entities (3)
-
communication state φ (exploration fraction)
no independent evidence
-
mechanism decomposition ratio δ = cD/(cQ+cD+cS)
no independent evidence
-
artifact-mediated buffering
no independent evidence
Cite this review
Pith. "Pith review of When Does Personality Composition Matter for Multi-Agent LLM Teams?." pith.science (2026). https://pith.science/paper/RV37EXFV
@misc{pith2026260627443,
author = {Pith},
title = {Pith review of: When Does Personality Composition Matter for Multi-Agent LLM Teams?},
year = {2026},
howpublished = {\url{https://pith.science/paper/RV37EXFV}},
note = {Machine review of arXiv:2606.27443}
}
read the original abstract
Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-explored. Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but the relationship between communication style and task performance has not been systematically examined across multiple domains. In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on three task domains: structured coding, open-ended research collaboration, and competitive bargaining. We find that personality effects depend critically on task structure. In coding tasks, low agreeableness leads to large communication shifts that have little effect on milestone completion. In open-ended collaboration and bargaining, the same manipulation substantially degrades performance. We discuss implications for multi-agent system design and the limits of personality manipulation.
Figures
Reference graph
Works this paper leans on
-
[1]
Accessed: 2026-03-
URL https: //www.anthropic.com/engineering/building-effective-agents/. Accessed: 2026-03-
2026
-
[2]
URL https://www.anthropic. com/news/claude-sonnet-4-5. Accessed: 2026-03-29. Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitu- tional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073,
Pith/arXiv arXiv 2026
-
[3]
Evalu- ating large language models trained on code.arXiv preprint arXiv:2107.03374,
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evalu- ating large language models trained on code.arXiv preprint arXiv:2107.03374,
-
[4]
Yifan Duan, Yihong Tang, Xuefeng Bai, Kehai Chen, Juntao Li, and Min Zhang. The power of personality: A human simulation perspective to investigate large language model agents.arXiv preprint arXiv:2502.20859,
-
[5]
How personality traits influence negotiation outcomes? a simulation based on large language models
Yin Jou Huang and Rafik Hadfi. How personality traits influence negotiation outcomes? a simulation based on large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 10336–10351,
2024
-
[6]
Gpt-4o system card.arXiv preprint arXiv:2410.21276,
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card.arXiv preprint arXiv:2410.21276,
-
[7]
Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770,
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770,
-
[8]
Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437,
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437,
-
[9]
Accessed: 2026-03-26
URL https://newsletter.pragmaticengineer.com/p/ ai-tooling-2026. Accessed: 2026-03-26. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 3...
2026
-
[10]
Keyu Pan and Yawen Zeng. Do llms possess a personality? making the mbti test an amazing evaluation for large language models.arXiv preprint arXiv:2307.16180,
-
[11]
Collab-overcooked: Benchmarking and evaluating large lan- guage models as collaborative agents
Haochen Sun, Shuwen Zhang, Lujie Niu, Lei Ren, Hao Xu, Hao Fu, Fangkun Zhao, Caixia Yuan, and Xiaojie Wang. Collab-overcooked: Benchmarking and evaluating large lan- guage models as collaborative agents. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 4922–4951,
2025
-
[12]
Sixiong Xie, Zhuofan Shi, Haiyang Shen, Gang Huang, Yun Ma, and Xiang Jing
Accessed: 2025-03-29. Sixiong Xie, Zhuofan Shi, Haiyang Shen, Gang Huang, Yun Ma, and Xiang Jing. M3-bench: Process-aware evaluation of llm agents social behaviors in mixed-motive games.arXiv preprint arXiv:2601.08462,
arXiv 2025
-
[13]
Personalllm: Tailoring llms to individual preferences.arXiv preprint arXiv:2409.20296,
Thomas P Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong. Personalllm: Tailoring llms to individual preferences.arXiv preprint arXiv:2409.20296,
This paper was first reviewed by grok-4.5 on July 15, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.