REVIEW 3 major objections 4 minor 3 cited by
LLM-Based Social Simulations Require a Boundary
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LLM-based social simulations need explicit boundaries in validation and claim scope because current models act as an 'average persona' with insufficient behavioral variance.
desk verdict A useful boundary checklist for LLM-based social simulation, slightly oversold by the 'fundamental' framing, but the practical recommendations hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the variance-mean framework, a two-dimensional diagnosis of simulation fidelity: the mean tells whether LLM output centers on the human average, and the variance tells whether the population of agents spreads like a real population. Its companion concept is the 'average persona,' the tendency of likelihood-trained models to concentrate on high-frequency responses and suppress tail behavior. The framework separates two failure cases—low variance with aligned mean, which permits qualitative collective-pattern claims but not quantitative or individual-trajectory claims, and low variance with deviated mean, which blocks most inferences about real societies. The same pairing, applied to 21 reviewed studies as a coding scheme (heterogeneity requirement, ground truth, mean check, variance check, claim level, sensitivity analysis), produces the paper's empirical observation that validation practice lags heterogeneity requirements.
What would settle it
Run the same battery of LLM social simulations—Keynesian Beauty Contest, public-opinion surveys, iterated prisoner's dilemma, and evacuation or market tasks—against the published human datasets used as benchmarks, measuring the variance ratio of simulated to human behavior; if simulated distributions are statistically indistinguishable from or wider than human ones on most tasks, the average-persona boundary claim would be refuted. A second check would have independent coders reapply the appendix criteria to the same 21 papers, since large disagreement on which papers check variance would undercut the review's empirical counts.
Extended reading notes
Core claim
The central claim is that LLM-based social simulations are boundary-limited, not merely imperfect: clear lines must be drawn around validation requirements and claim levels before a simulation can contribute meaningfully to social science. The paper asserts that current LLMs act as an 'average persona'—their training objective rewards high-frequency, mainstream responses, so a population of agents produces too little behavioral variance even when its average behavior aligns with human data. In the variance-mean framework, mean-aligned low variance supports only qualitative collective-pattern claims, while mean-deviated low variance makes a simulation largely inapplicable to real societies. A review of 21 studies from 2023 to 2025 finds that all ground-truth-checking papers assess mean alignment, but fewer than half assess variance, that most variance checks report lower-than-human diversity, and that this validation depth often falls below what high-heterogeneity research questions demand. The paper's positive position is that boundary-aware simulation—where validation depth tracks the heterogeneity requirement of the question and claims are scoped accordingly—can still contribute genuine social-science insight.
Load-bearing premise
The argument depends on the empirical premise that current LLM agents genuinely generate narrower behavioral distributions than human populations across the social tasks under study, rather than merely doing so in the examples reviewed; if prompting, scaling, or sampling can close that variance gap, the average-persona boundary would become a usage problem, not an inherent limit.
Editorial extensions
If this is right
- Researchers running LLM simulations on questions about polarization, tipping points, inequality, or minority influence should treat a mean-alignment check as insufficient and should measure the variance of simulated behavior against human baselines before drawing conclusions.
- Publications in this area will need to report variance statistics alongside central tendency, and when variance is low, phrase findings as qualitative collective patterns rather than exact frequencies, shares, or individual trajectories.
- Claims about precise quantitative matches, such as a reported accuracy percentage against ground truth, or about explanatory individual behaviors, will be recognized as outside the boundary for low-variance simulations.
- Domains with accumulated human experimental data, such as game theory, market experiments, and public-opinion surveys, become the preferred test beds for LLM simulation because variance can be benchmarked, while hard-to-validate domains like large-scale network dynamics will need shared benchmarks first.
- If the field follows the paper's recommendations, the debate shifts from whether LLMs are valid at all to under which heterogeneity conditions and claim levels a particular simulation is valid.
Reading between the lines
- A direct extension of this reasoning is that the variance-mean check becomes a general acceptance criterion for any generative-agent simulation, not just current LLMs; future models should face the same benchmark.
- A testable next step the authors do not develop is a standardized variance-audit suite built from existing human experiments, so each new model generation can be tracked against the average-persona gap.
- The paper's logic implies that efforts to increase diversity through persona prompts, role-play, or multi-agent setups should be judged by output variance against real distributions, since input heterogeneity does not guarantee output heterogeneity.
- If the average-persona gap is rooted in likelihood training, then decoding changes such as temperature, sampling, and contrastive decoding, as well as fine-tuning on tail-heavy data, become plausible interventions; measuring their effect on output variance would show whether the boundary is architectural or tunable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that LLM-based social simulations require explicit boundaries on validation and claim levels. The authors introduce a mean-variance framework, contending that LLMs act as an 'average persona' with systematically low behavioral variance, which limits their usefulness for simulating complex social dynamics. They support this with a training-objective argument, selected empirical examples, and a hand-coded review of 21 representative studies. The review reports that 14 of 21 papers include ground-truth comparisons, but only 9 of these explicitly assess variance, and most variance checks find lower variance than human baselines. The paper recommends matching validation depth to the heterogeneity requirements of the research question, reporting variance alongside mean alignment, and constraining claims to collective-level qualitative patterns when variance is insufficient.
Significance. If correct, the paper's core message is valuable: it reframes the debate from whether LLMs can simulate humans to what can validly be claimed from mean-aligned, low-variance simulations. The proposed checklist is concrete and actionable, and the paper engages both optimistic and skeptical literatures in a balanced way. The appendix materials, the decision heuristic for heterogeneity requirements, and the 'contradiction audit' suggestion are useful contributions. However, the empirical foundation is weaker than the conceptual framework: the review coding is subjective and not independently verifiable, and the 'fundamental' limitation claim is asserted more strongly than the evidence warrants. The central insight is likely right, but the manuscript would benefit from tempering its claims and adding transparency to the review.
major comments (3)
- [Abstract; §4.1; §4.2] The abstract and §4.1 state that the average persona 'fundamentally limits' LLMs' ability to capture behavioral diversity, and Contribution (1) refers to 'inherent limitations that fundamentally determine their reliability.' This premise is load-bearing for the paper's central claim-boundary recommendation, yet the manuscript never tests whether output variance can be restored by decoding choices (e.g., temperature, top-p), persona richness, or population mixtures. The training-objective argument (likelihood maximization rewards common responses) explains a tendency, not an immovable limit. The evidence cited in §4.2 shows that current heterogeneity-enhancement methods 'often fail,' which is a contingent empirical finding. If some configuration reproduces both human mean and variance, the recommendation to constrain claims to collective-qualitative patterns would be too strong as a general rule. I recommend either softening 'fundamentally' to 'currently tends to' while explicitly scoping the claim to current models and default settings, or adding a controlled comparison of LLM variance against a human benchmark across sampling and persona conditions.
- [§5, Table 1, Appendix C] The systematic review is a central contribution (Contribution 3), and its headline counts (14/21 ground truth; 9/14 variance checked) depend on judgment calls that the authors acknowledge in Appendix C.1. No inter-coder reliability, coding sheet, or raw extracted data are provided, so a different reviewer could code borderline cases differently. For example, the assignment of heterogeneity requirements (High/Medium/Low) and the classification of 'Mean' and 'Var.' results (Aligned/Deviated/Mixed/Low) are not tied to explicit thresholds or verbatim excerpts. I ask the authors to make the coding transparent: include per-paper excerpts or a detailed coding table with the rationale for each dimension, and report inter-coder agreement if more than one coder was involved. Without this, the empirical support for the 'validation gap' claim is not independently verifiable, even though the conceptual argument may still stand.
- [§4.1; §5.3, Recommendation 2] The mean-variance framework is intuitive but under-specified. 'Variance' is defined only as 'the diversity and spread of behaviors' (§4.1), and no concrete metric is proposed for measuring it or for comparing against human baselines. This matters because Recommendation 2 asks researchers to 'report variance explicitly' and Recommendation 1 asks them to match validation depth to heterogeneity requirements; without an operationalization, these are hard to implement consistently across studies. I suggest the authors specify at least one example metric (e.g., variance or entropy of a target action distribution, inter-agent behavioral distance) and state how to determine whether variance is 'sufficient' relative to a human reference distribution, or explicitly delegate this to future work with a concrete proposal.
minor comments (4)
- [Abstract; §4.1] There are typographical issues: 'argues thatLLM-based' appears in the abstract and in §4.1, and 'Reynolds (1987)' should be 'Reynolds (1987)'s.'
- [Table 1; Appendix C.3] The Table 1 legend defines Aligned/Deviated/Mixed for Mean/Variance results but does not explain the entry 'Low' in the Var. column; Appendix C.3 defines 'Lower' for variance comparisons against ground truth, so the table's 'Low' should be either 'Lower' or explicitly defined.
- [Abstract; §5.1] The abstract's 'fewer than half explicitly assess behavioral variance' refers to the full sample (9/21), while §5.1 reports 'fewer explicitly assessed variance (9 of 14)' among ground-truth papers; consider clarifying the denominator in both places to avoid misreading.
- [References] The reference list contains duplicate entries for Hua et al. (2023/2024) and Fontana et al. (2024/2025); consider merging or cross-referencing to avoid confusion.
Circularity Check
No significant circularity: the boundary argument is conceptual and empirically grounded; self-citations are illustrative, not load-bearing.
full rationale
This is a position paper whose central claim—that LLM-based social simulations require clear boundaries and that LLMs' 'average persona' tendency limits behavioral diversity—is argued conceptually and supported by a systematic review, not derived by fitting parameters or by defining the conclusion into the premises. The average-persona phenomenon is supported by multiple independent external studies (e.g., Bisbee et al. 2024, Argyle et al. 2023, Xie et al. 2024, del Rio-Chanona et al. 2025), so the authors' citation of their own prior works (Wu et al. 2023, 2024) as illustrative examples is not load-bearing. The review's coding criteria involve acknowledged judgment calls, but that is a reliability limitation, not circularity: the counts do not reduce by construction to the paper's conclusions. The claim that temperature, sampling, or persona richness might restore variance is a substantive challenge to the strength of the 'fundamental limits' assertion, but it is a correctness concern, not a circular-derivation concern. No equation or fitted value is renamed as a prediction, and no self-cited uniqueness theorem forces the authors' choice. The paper is self-contained as an argumentative and empirical review, so no circularity is present.
Assumptions & free parameters
assumptions (3)
- domain assumption The primary objective of social simulation is explaining social patterns and generating hypotheses, not replication or prediction.
- domain assumption Research questions whose answers depend on distributional properties require higher agent heterogeneity.
- domain assumption Human population variance is the appropriate ground truth for judging LLM behavioral variance.
Cite this review
Pith. "Pith review of LLM-Based Social Simulations Require a Boundary." pith.science (2026). https://pith.science/paper/RFPCXDWY
@misc{pith2026250619806,
author = {Pith},
title = {Pith review of: LLM-Based Social Simulations Require a Boundary},
year = {2026},
howpublished = {\url{https://pith.science/paper/RFPCXDWY}},
note = {Machine review of arXiv:2506.19806}
}
read the original abstract
This position paper argues that LLM-based social simulations require clear boundaries to make meaningful contributions to social science. While Large Language Models (LLMs) offer promising capabilities for simulating human behavior, their tendency to produce homogeneous outputs, acting as an "average persona", fundamentally limits their ability to capture the behavioral diversity essential for complex social dynamics. We examine why heterogeneity matters for social simulations and how current LLMs fall short, analyzing the relationship between mean alignment and variance in LLM-generated behaviors. Through a systematic review of representative studies, we find that validation practices often fail to match the heterogeneity requirements of research questions: while most papers include ground truth comparisons, fewer than half explicitly assess behavioral variance, and most that do report lower variance than human populations. We propose that researchers should: (1) match validation depth to the heterogeneity demands of their research questions, (2) explicitly report variance alongside mean alignment, and (3) constrain claims to collective-level qualitative patterns when variance is insufficient. Rather than dismissing LLM-based simulation, we advocate for a boundary-aware approach that ensures these methods contribute genuine insights to social science.
Figures
Forward citations
Cited by 3 Pith papers
-
Will Scaling Improve Social Simulation with LLMs?
Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.
-
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
Grounding LLM agents with ACS demographics and ATUS time-use routines substantially improved reproduction of observed activity profiles under normal and heatwave conditions in Philadelphia, though over half of the hea...
-
Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community
Clustering of Moltbook submolt descriptions shows agent-created communities organize into human-mimetic, silicon-centric, and proto-economic themes, but the categories were partly prescribed by the analysis prompt.
Reference graph
Works this paper leans on
-
[1]
C.1. Research Question Type and Heterogeneity Requirement Our classification draws on established frameworks in computational social science and ABM literature. Following Squazzoni et al. (2014) and Edmonds et al. (2019), we recognize that different modeling purposes impose different requirements on agent heterogeneity. We adapt this idea to categorize re...
work page 2014
-
[3]
Bernardelle, P., Fr¨ohling, L., Civelli, S., Lunardi, R., Roitero, K., and Demartini, G. Mapping and influencing the po- litical ideology of large language models using synthetic personas.arXiv preprint arXiv:2412.14843,
-
[6]
Chen, R., Li, Y ., Yang, J., Feng, Y ., Zhou, J. T., Wu, J., and Liu, Z. Identifying and mitigating social bias knowledge in language models. InFindings of the Association for Com- putational Linguistics: NAACL 2025, pp. 651–672,
work page 2025
-
[7]
Compost: Charac- terizing and evaluating caricature in llm simulations
Cheng, M., Piccardi, T., and Yang, D. Compost: Charac- terizing and evaluating caricature in llm simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 10853–10875,
work page 2023
-
[9]
Simulating opinion dynamics with networks of llm-based agents
Chuang, Y .-S., Goyal, A., Harlalka, N., Suresh, S., Hawkins, R., Yang, S., Shah, D., Hu, J., and Rogers, T. Simulating opinion dynamics with networks of llm-based agents. In Findings of the Association for Computational Linguistics: NAACL 2024, pp. 3326–3346, 2024a. Chuang, Y .-S., Nirunwiroj, K., Studdiford, Z., Goyal, A., Frigo, V ., Yang, S., Shah, D....
work page 2024
-
[11]
Behavioral and Topological Heterogeneities in Network Versions of Schelling's Segregation Model
Deter, W. and Sayama, H. Behavioral and topological hetero- geneities in network versions of schelling’s segregation model.arXiv preprint arXiv:2408.05623,
-
[12]
Fontana, N., Pierri, F., and Aiello, L. M. Nicer than humans: How do large language models behave in the prisoner’s dilemma?arXiv preprint arXiv:2406.13605,
-
[16]
11 LLM-Based Social Simulations Require a Boundary Gandhi, K., Sadigh, D., and Goodman, N. D. Strategic reasoning with language models.arXiv preprint arXiv:2305.19165,
Show all 65 references
-
[17]
S3: Social-network simulation system with large language model-empowered agents.arXiv preprint arXiv:2307.14984,
Gao, C., Lan, X., Lu, Z., Mao, J., Piao, J., Wang, H., Jin, D., and Li, Y . S3: Social-network simulation system with large language model-empowered agents.arXiv preprint arXiv:2307.14984,
-
[18]
Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,
Ge, T., Chan, X., Wang, X., Yu, D., Mi, H., and Yu, D. Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,
-
[19]
and Toubia, O
Gui, G. and Toubia, O. The challenge of using llms to simulate human behavior: A causal inference perspective. arXiv preprint arXiv:2312.15524,
-
[20]
Can large language models play games? a case study of a self-play approach
Guo, H., Liu, Z., Zhang, Y ., and Wang, Z. Can large language models play games? a case study of a self-play approach. arXiv preprint arXiv:2403.05632,
-
[21]
”guinea pig trials” utilizing gpt: A novel smart agent-based modeling approach for studying firm competition and collusion.arXiv preprint arXiv:2308.10974,
Han, X., Wu, Z., and Xiao, C. ”guinea pig trials” utilizing gpt: A novel smart agent-based modeling approach for studying firm competition and collusion.arXiv preprint arXiv:2308.10974,
-
[22]
M., Antunes, L., Pav´on Mestras, J., et al
Hassan, S., Arroyo, J., Gal ´an Ordax, J. M., Antunes, L., Pav´on Mestras, J., et al. Asking the oracle: Introducing forecasting principles into agent-based modelling. Journal of artificial societies and social simulation. 2013, V . 16, n. 3,
2013
-
[23]
J., and Xie, X
Hu, Z., Lian, J., Xiao, Z., Xiong, M., Lei, Y ., Wang, T., Ding, K., Xiao, Z., Yuan, N. J., and Xie, X. Population-aligned persona generation for llm-based social simulation.arXiv preprint arXiv:2509.10127,
-
[25]
War and peace (waragent): Llm-based multi-agent simulation of world wars.arXiv preprint arXiv:2311.17227,
Hua, W., Fan, L., Li, L., Mei, K., Ge, Y ., Hemphill, L., Zhang, Y ., et al. War and peace (waragent): Llm-based multi-agent simulation of world wars.arXiv preprint arXiv:2311.17227,
-
[26]
Social science meets llms: How reliable are large language models in so- cial simulations?arXiv preprint arXiv:2410.23426,
Huang, Y ., Yuan, Z., Zhou, Y ., Guo, K., Wang, X., Zhuang, H., Sun, W., Sun, L., Wang, J., Ye, Y ., et al. Social science meets llms: How reliable are large language models in so- cial simulations?arXiv preprint arXiv:2410.23426,
-
[27]
Understanding llm agent behaviours via game theory: Strategy recognition, biases and multi-agent dynamics.arXiv preprint arXiv:2512.07462,
Huynh, T.-K., Dao-Sy, D.-M., Cao, T.-B., Le, P.-H., Nguyen, H.-D., Nguyen-Lam, P.-Q., Nguyen-V o, M.-L., Pham, H.-P., Pham, P.-H., Than, T.-K., et al. Understanding llm agent behaviours via game theory: Strategy recognition, biases and multi-agent dynamics.arXiv preprint arXiv...
-
[29]
and T ¨ornberg, P
Larooij, M. and T ¨ornberg, P. Do large language models solve the problems of agent-based modeling? a critical review of generative social simulations.arXiv preprint arXiv:2504.03274,
-
[30]
Llm generated persona is a promise with a catch.arXiv preprint arXiv:2503.16527, 2025a
Li, A., Chen, H., Namkoong, H., and Peng, T. Llm generated persona is a promise with a catch.arXiv preprint arXiv:2503.16527, 2025a. Li, C. J., Wu, J., Mo, Z., Qu, A., Tang, Y ., Zhao, K. I., Gan, Y ., Fan, J., Yu, J., Zhao, J., et al. Simulating society requires simulating th...
-
[31]
Diversedia- logue: A methodology for designing chatbots with human- like diversity.arXiv preprint arXiv:2409.00262,
Lin, X., Yu, X., Aich, A., Giorgi, S., and Ungar, L. Diversedia- logue: A methodology for designing chatbots with human- like diversity.arXiv preprint arXiv:2409.00262,
-
[32]
Evaluating large language model biases in persona-steered generation
Liu, A., Diab, M., and Fried, D. Evaluating large language model biases in persona-steered generation. InFindings of the Association for Computational Linguistics ACL 2024, pp. 9832–9850, 2024a. Liu, X., Yu, H., Zhang, H., Xu, Y ., Lei, X., Lai, H., Gu, Y ., Ding, H., Men, K.,...
2024 arXiv
-
[33]
How does the heterogeneity of members affect the evolution of group opinions?Discrete Dynamics in Nature and Society, 2021(1):8827048,
Lu, A., Ling, H., and Ding, Z. How does the heterogeneity of members affect the evolution of group opinions?Discrete Dynamics in Nature and Society, 2021(1):8827048,
2021
-
[35]
Computational experiments meet large language model based agents: A survey and perspective.arXiv preprint arXiv:2402.00262,
Ma, Q., Xue, X., Zhou, D., Yu, X., Liu, D., Zhang, X., Zhao, Z., Shen, Y ., Ji, P., Li, J., et al. Computational experiments meet large language model based agents: A survey and perspective.arXiv preprint arXiv:2402.00262,
-
[36]
J., and Ohno-Machado, L
Ma, X., Zhu, R., Wang, Z., Xiong, J., Chen, Q., Tang, H., Camp, L. J., and Ohno-Machado, L. Enhancing patient- centric communication: Leveraging llms to simulate patient perspectives.arXiv preprint arXiv:2501.06964,
-
[37]
Towards a holistic landscape of situated theory of mind in large language models.arXiv preprint arXiv:2310.19619,
Ma, Z., Sansom, J., Peng, R., and Chai, J. Towards a holistic landscape of situated theory of mind in large language models.arXiv preprint arXiv:2310.19619,
-
[40]
and Xu, W
Naous, T. and Xu, W. On the origin of cultural biases in language models: From pre-training data to linguistic phenomena.arXiv preprint arXiv:2501.04662,
-
[41]
Beyond demographics: Fine-tuning large language models to predict individuals’ subjective text perceptions.arXiv preprint arXiv:2502.20897,
Orlikowski, M., Pei, J., R¨ottger, P., Cimiano, P., Jurgens, D., and Hovy, D. Beyond demographics: Fine-tuning large language models to predict individuals’ subjective text perceptions.arXiv preprint arXiv:2502.20897,
-
[42]
S., Zou, C
Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Willer, R., Liang, P., and Bernstein, M. S. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109,
-
[45]
P., Galley, M., Caruana, R., and Gao, J
Singh, C., Inala, J. P., Galley, M., Caruana, R., and Gao, J. Rethinking interpretability in the era of large language models.arXiv preprint arXiv:2402.01761,
-
[46]
R., Jain, R., and Mehta, S
Surve, A., Rathod, A., Surana, M., Malpani, G., Shamraj, A., Sankepally, S. R., Jain, R., and Mehta, S. S. Mul- tiagent simulators for social networks.arXiv preprint arXiv:2311.14712,
-
[47]
Gensim: A general social simulation platform with large language model based agents
Tang, J., Gao, H., Pan, X., Wang, L., Tan, H., Gao, D., Chen, Y ., Chen, X., Lin, Y ., Li, Y ., et al. Gensim: A general social simulation platform with large language model based agents. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ...
2025
-
[48]
Systematic biases in llm simulations of debates
Taubenfeld, A., Dover, Y ., Reichart, R., and Goldstein, A. Systematic biases in llm simulations of debates. InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 251–267,
2024
-
[50]
Large language models as urban residents: An llm agent framework for personal mobility generation.Advances in Neural Information Processing Systems, 37:124547–124574, 2024a
Wang, J., Jiang, R., Yang, C., Wu, Z., Onizuka, M., Shibasaki, R., Koshizuka, N., Xiao, C., et al. Large language models as urban residents: An llm agent framework for personal mobility generation.Advances in Neural Information Processing Systems, 37:124547–124574, 2024a. 15 L...
-
[51]
From chatgpt to deepseek: Can llms simulate humanity?arXiv preprint arXiv:2502.18210, 2025c
Wang, Q., Tang, Z., and He, B. From chatgpt to deepseek: Can llms simulate humanity?arXiv preprint arXiv:2502.18210, 2025c. Wang, Q., Wu, J., Tang, Z., Luo, B., Chen, N., Chen, W., and He, B. What limits llm-based human simulation: Llms or our design?arXiv preprint arXiv:2501....
-
[52]
Smart agent-based modeling: On the use of large language models in computer simulations.arXiv preprint arXiv:2311.06330,
Wu, Z., Peng, R., Han, X., Zheng, S., Zhang, Y ., and Xiao, C. Smart agent-based modeling: On the use of large language models in computer simulations.arXiv preprint arXiv:2311.06330,
-
[53]
Shall we team up: Exploring spontaneous cooperation of competing llm agents
Wu, Z., Peng, R., Zheng, S., Liu, Q., Han, X., Kwon, B., Onizuka, M., Tang, S., and Xiao, C. Shall we team up: Exploring spontaneous cooperation of competing llm agents. InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 5163–5186,
2024
-
[55]
Evaluating and enhancing llms agent based on theory of mind in guandan: A multi-player cooperative game under imperfect information.arXiv preprint arXiv:2408.02559,
Yim, Y ., Chan, C., Shi, T., Deng, Z., Fan, W., Zheng, T., and Song, Y . Evaluating and enhancing llms agent based on theory of mind in guandan: A multi-player cooperative game under imperfect information.arXiv preprint arXiv:2408.02559,
-
[56]
Exploring collaboration mechanisms for llm agents: A social psychology view.arXiv preprint arXiv:2310.02124,
Zhang, J., Xu, X., Zhang, N., Liu, R., Hooi, B., and Deng, S. Exploring collaboration mechanisms for llm agents: A social psychology view.arXiv preprint arXiv:2310.02124,
-
[57]
Don’t go to extremes: Revealing the excessive sensitivity and calibration limitations of llms in implicit hate speech detection
Zhang, M., He, J., Ji, T., and Lu, C.-T. Don’t go to extremes: Revealing the excessive sensitivity and calibration limitations of llms in implicit hate speech detection. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
-
[58]
K-level reasoning with large language models
Zhang, Y ., Mao, S., Ge, T., Wang, X., Xia, Y ., Lan, M., and Wei, F. K-level reasoning with large language models. arXiv e-prints, pp. arXiv–2402, 2024b. Zhang, Z., Hu, F., Lee, J., Shi, F., Kordjamshidi, P., Chai, J., and Ma, Z. Do vision-language models represent space and ...
-
[59]
H., Wang, X., Zhu, H., Wang, W., and Sap, M
Zhou, J., Huang, J.-t., Zhou, X., Lam, M. H., Wang, X., Zhu, H., Wang, W., and Sap, M. The pimmur principles: Ensuring validity in collective behavior of llm societies. arXiv preprint arXiv:2509.18052,
-
[60]
Sotopia: Interactive evaluation for social intelligence in language agents.arXiv preprint arXiv:2310.11667,
Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y ., Fried, D., Neubig, G., et al. Sotopia: Interactive evaluation for social intelligence in language agents.arXiv preprint arXiv:2310.11667,
-
[61]
apophenia
17 LLM-Based Social Simulations Require a Boundary A. Related Works A.1. Computational Social Science Social phenomena typically arise from the interactions of intelligent, adaptive agents under dynamic conditions (Eidelson, 1997; San Miguel et al., 2012). Even when we fully u...
1997
-
[63]
(3) Exploratory studies have demonstrated human-like behavior, with performance approaching that of humans in certain experiments (Anthis et al., 2025)
and improving the fidelity of complex behaviors such as interaction, collaboration, and gaming (Ma et al., 2024). (3) Exploratory studies have demonstrated human-like behavior, with performance approaching that of humans in certain experiments (Anthis et al., 2025). However, r...
2024
-
[65]
Would the research question still be answerable if all agents behaved identically at the population mean?
• Low: Questions primarily concerning equilibrium existence or central tendency dynamics, where the specific distribution shape matters less. 19 LLM-Based Social Simulations Require a Boundary •Medium: Questions involving some distributional aspects but where central tendencie...
2024
-
[100]
Can markets clear?
C.3. Alignment Assessment Mean Alignmentwas marked as checked (✓) if the paper explicitly compared average LLM agent behavior against human baselines. Results were categorized as: •Aligned: LLM mean behavior is statistically or qualitatively consistent with human average. •Dev...
2025
-
[1971]
Personality traits in large language models.arXiv preprint arXiv:2307.00184,
Serapio-Garc´ıa, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., and Matari´c, M. Personality traits in large language models.arXiv preprint arXiv:2307.00184,
-
[1988]
United in diversity? contextual biases in llm-based predictions of the 2024 european parliament elections.arXiv preprint arXiv:2409.09045,
von der Heyde, L., Haensch, A.-C., and Wenz, A. United in diversity? contextual biases in llm-based predictions of the 2024 european parliament elections.arXiv preprint arXiv:2409.09045,
2024 arXiv
-
[1996]
K., Bhatia, S., and Chakraborty, T
Chatterjee, A., Renduchintala, H. K., Bhatia, S., and Chakraborty, T. Posix: A prompt sensitivity index for large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 14550–14565,
2024
-
[1999]
and Schelling’s segregation model (Schelling, 1971), which illustrate how wealth gaps or segregation patterns can emerge from individual interactions. Despite disagreements and inconsistencies within social science theories, many works agree that social interaction is the fund...
1971
-
[2005]
SIS 2005., pp. 201–208. IEEE,
2005
-
[2006]
Modeling and math- ematical analysis of swarms of microscopic robots
Galstyan, A., Hogg, T., and Lerman, K. Modeling and math- ematical analysis of swarms of microscopic robots. In Proceedings 2005 IEEE Swarm Intelligence Symposium,
2005
-
[2009]
Explaining large language models decisions using shapley values.arXiv preprint arXiv:2404.01332,
Mohammadi, B. Explaining large language models decisions using shapley values.arXiv preprint arXiv:2404.01332,
-
[2010]
Emergence of social norms in generative agent societies: principles and architecture.arXiv preprint arXiv:2403.08251,
14 LLM-Based Social Simulations Require a Boundary Ren, S., Cui, Z., Song, R., Wang, Z., and Hu, S. Emergence of social norms in generative agent societies: principles and architecture.arXiv preprint arXiv:2403.08251,
-
[2011]
Specializing large language models to simulate survey response distributions for global populations.arXiv preprint arXiv:2502.07068,
10 LLM-Based Social Simulations Require a Boundary Cao, Y ., Liu, H., Arora, A., Augenstein, I., R¨ottger, P., and Hershcovich, D. Specializing large language models to simulate survey response distributions for global populations.arXiv preprint arXiv:2502.07068,
-
[2014]
Are large language models (llms) good social predictors? InFindings of the Association for Computational Linguistics: EMNLP 2024, pp
Yang, K., Li, H., Wen, H., Peng, T.-Q., Tang, J., and Liu, H. Are large language models (llms) good social predictors? InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 2718–2730, 2024a. Yang, Y ., Duan, H., Liu, J., and Tam, K. Y . Llm-measure: Gene...
2024 arXiv
-
[2015]
M., Pangallo, M., and Hommes, C
del Rio-Chanona, R. M., Pangallo, M., and Hommes, C. Can generative ai agents behave like humans? evidence from laboratory market experiments.arXiv preprint arXiv:2505.07457,
-
[2017]
G., Ortu, F., Strausz, A., Sachan, M., Mihalcea, R., et al
Jin, Z., Kleiman-Weiner, M., Piatti, G., Levine, S., Liu, J., Adauto, F. G., Ortu, F., Strausz, A., Sachan, M., Mihalcea, R., et al. Multilingual trolley problems for language models. InPluralistic Alignment Workshop at NeurIPS 2024,
2024
-
[2018]
R., Liu, R., Richardson, S
Anthis, J. R., Liu, R., Richardson, S. M., Kozlowski, A. C., Koch, B., Evans, J., Brynjolfsson, E., and Bernstein, M. Llm social simulations are a promising research method. arXiv preprint arXiv:2504.02234,
-
[2021]
Lu, S. E. Strategic interactions between large language models-based agents in beauty contests.arXiv preprint arXiv:2404.08492,
-
[2022]
From individual to society: A survey on social simulation driven by large language model-based agents.arXiv preprint arXiv:2412.03563,
Mou, X., Ding, X., He, Q., Wang, L., Liang, J., Zhang, X., Sun, L., Lin, J., Zhou, J., Huang, X., et al. From individual to society: A survey on social simulation driven by large language model-based agents.arXiv preprint arXiv:2412.03563,
-
[2023]
Observing micromotives and macrobehavior of large language models.arXiv preprint arXiv:2412.10428,
Cheng, Y ., Qu, X., Goldsack, T., Lin, C., and Chen, C.-C. Observing micromotives and macrobehavior of large language models.arXiv preprint arXiv:2412.10428,
-
[2024]
Can we count on llms? the fixed-effect fallacy and claims of gpt-4 capabilities.arXiv preprint arXiv:2409.07638,
Ball, T., Chen, S., and Herley, C. Can we count on llms? the fixed-effect fallacy and claims of gpt-4 capabilities.arXiv preprint arXiv:2409.07638,
-
[2025]
Anomalous fluctua- tions in minority games and related multi-agent models of financial markets.arXiv preprint physics/0608091,
Galla, T., Mosetti, G., and Zhang, Y .-C. Anomalous fluctua- tions in minority games and related multi-agent models of financial markets.arXiv preprint physics/0608091,
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.