Pith. sign in

REVIEW 3 major objections 4 minor 3 cited by

LLM-Based Social Simulations Require a Boundary

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read LLM-based social simulations need explicit boundaries in validation and claim scope because current models act as an 'average persona' with insufficient behavioral variance.

desk verdict A useful boundary checklist for LLM-based social simulation, slightly oversold by the 'fundamental' framing, but the practical recommendations hold. read the letter →

arxiv 2506.19806 v3 pith:RFPCXDWY submitted 2025-06-24 cs.CY cs.CLcs.MA

classification cs.CYcs.CLcs.MA
keywords LLM-basedsocialsimulationaveragepersonabehavioralvarianceagentheterogeneityvalidationboundariesclaimlevelsmeanalignmentscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that LLM-based social simulations of human societies cannot be treated as unconstrained stand-ins for real populations: researchers need explicit boundaries on what they validate and on what they claim to know. Its core negative finding is the 'average persona' phenomenon, where LLM agents produce homogeneous, low-variance behavior that matches human averages but suppresses the diversity actual populations exhibit. A systematic review of 21 representative studies shows that most check whether simulated behavior matches human means, but fewer than half explicitly check behavioral variance, and most that do find LLM variance below human variance. The paper therefore proposes a variance-mean framework: match validation depth to how much heterogeneity the research question requires, report variance alongside mean alignment, and restrict claims to qualitative collective patterns when variance is insufficient. If correct, the field overestimates the social-science value of many simulations whose means look right but whose distributions are too narrow to support conclusions about tipping points, minority influence, or distributional outcomes.

What carries the argument

The load-bearing device is the variance-mean framework, a two-dimensional diagnosis of simulation fidelity: the mean tells whether LLM output centers on the human average, and the variance tells whether the population of agents spreads like a real population. Its companion concept is the 'average persona,' the tendency of likelihood-trained models to concentrate on high-frequency responses and suppress tail behavior. The framework separates two failure cases—low variance with aligned mean, which permits qualitative collective-pattern claims but not quantitative or individual-trajectory claims, and low variance with deviated mean, which blocks most inferences about real societies. The same pairing, applied to 21 reviewed studies as a coding scheme (heterogeneity requirement, ground truth, mean check, variance check, claim level, sensitivity analysis), produces the paper's empirical observation that validation practice lags heterogeneity requirements.

What would settle it

Run the same battery of LLM social simulations—Keynesian Beauty Contest, public-opinion surveys, iterated prisoner's dilemma, and evacuation or market tasks—against the published human datasets used as benchmarks, measuring the variance ratio of simulated to human behavior; if simulated distributions are statistically indistinguishable from or wider than human ones on most tasks, the average-persona boundary claim would be refuted. A second check would have independent coders reapply the appendix criteria to the same 21 papers, since large disagreement on which papers check variance would undercut the review's empirical counts.

Watch

Extended reading notes

Core claim

The central claim is that LLM-based social simulations are boundary-limited, not merely imperfect: clear lines must be drawn around validation requirements and claim levels before a simulation can contribute meaningfully to social science. The paper asserts that current LLMs act as an 'average persona'—their training objective rewards high-frequency, mainstream responses, so a population of agents produces too little behavioral variance even when its average behavior aligns with human data. In the variance-mean framework, mean-aligned low variance supports only qualitative collective-pattern claims, while mean-deviated low variance makes a simulation largely inapplicable to real societies. A review of 21 studies from 2023 to 2025 finds that all ground-truth-checking papers assess mean alignment, but fewer than half assess variance, that most variance checks report lower-than-human diversity, and that this validation depth often falls below what high-heterogeneity research questions demand. The paper's positive position is that boundary-aware simulation—where validation depth tracks the heterogeneity requirement of the question and claims are scoped accordingly—can still contribute genuine social-science insight.

Load-bearing premise

The argument depends on the empirical premise that current LLM agents genuinely generate narrower behavioral distributions than human populations across the social tasks under study, rather than merely doing so in the examples reviewed; if prompting, scaling, or sampling can close that variance gap, the average-persona boundary would become a usage problem, not an inherent limit.

Editorial extensions

If this is right

  • Researchers running LLM simulations on questions about polarization, tipping points, inequality, or minority influence should treat a mean-alignment check as insufficient and should measure the variance of simulated behavior against human baselines before drawing conclusions.
  • Publications in this area will need to report variance statistics alongside central tendency, and when variance is low, phrase findings as qualitative collective patterns rather than exact frequencies, shares, or individual trajectories.
  • Claims about precise quantitative matches, such as a reported accuracy percentage against ground truth, or about explanatory individual behaviors, will be recognized as outside the boundary for low-variance simulations.
  • Domains with accumulated human experimental data, such as game theory, market experiments, and public-opinion surveys, become the preferred test beds for LLM simulation because variance can be benchmarked, while hard-to-validate domains like large-scale network dynamics will need shared benchmarks first.
  • If the field follows the paper's recommendations, the debate shifts from whether LLMs are valid at all to under which heterogeneity conditions and claim levels a particular simulation is valid.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension of this reasoning is that the variance-mean check becomes a general acceptance criterion for any generative-agent simulation, not just current LLMs; future models should face the same benchmark.
  • A testable next step the authors do not develop is a standardized variance-audit suite built from existing human experiments, so each new model generation can be tracked against the average-persona gap.
  • The paper's logic implies that efforts to increase diversity through persona prompts, role-play, or multi-agent setups should be judged by output variance against real distributions, since input heterogeneity does not guarantee output heterogeneity.
  • If the average-persona gap is rooted in likelihood training, then decoding changes such as temperature, sampling, and contrastive decoding, as well as fine-tuning on tail-heavy data, become plausible interventions; measuring their effect on output variance would show whether the boundary is architectural or tunable.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position paper argues that LLM-based social simulations require explicit boundaries on validation and claim levels. The authors introduce a mean-variance framework, contending that LLMs act as an 'average persona' with systematically low behavioral variance, which limits their usefulness for simulating complex social dynamics. They support this with a training-objective argument, selected empirical examples, and a hand-coded review of 21 representative studies. The review reports that 14 of 21 papers include ground-truth comparisons, but only 9 of these explicitly assess variance, and most variance checks find lower variance than human baselines. The paper recommends matching validation depth to the heterogeneity requirements of the research question, reporting variance alongside mean alignment, and constraining claims to collective-level qualitative patterns when variance is insufficient.

Significance. If correct, the paper's core message is valuable: it reframes the debate from whether LLMs can simulate humans to what can validly be claimed from mean-aligned, low-variance simulations. The proposed checklist is concrete and actionable, and the paper engages both optimistic and skeptical literatures in a balanced way. The appendix materials, the decision heuristic for heterogeneity requirements, and the 'contradiction audit' suggestion are useful contributions. However, the empirical foundation is weaker than the conceptual framework: the review coding is subjective and not independently verifiable, and the 'fundamental' limitation claim is asserted more strongly than the evidence warrants. The central insight is likely right, but the manuscript would benefit from tempering its claims and adding transparency to the review.

major comments (3)
  1. [Abstract; §4.1; §4.2] The abstract and §4.1 state that the average persona 'fundamentally limits' LLMs' ability to capture behavioral diversity, and Contribution (1) refers to 'inherent limitations that fundamentally determine their reliability.' This premise is load-bearing for the paper's central claim-boundary recommendation, yet the manuscript never tests whether output variance can be restored by decoding choices (e.g., temperature, top-p), persona richness, or population mixtures. The training-objective argument (likelihood maximization rewards common responses) explains a tendency, not an immovable limit. The evidence cited in §4.2 shows that current heterogeneity-enhancement methods 'often fail,' which is a contingent empirical finding. If some configuration reproduces both human mean and variance, the recommendation to constrain claims to collective-qualitative patterns would be too strong as a general rule. I recommend either softening 'fundamentally' to 'currently tends to' while explicitly scoping the claim to current models and default settings, or adding a controlled comparison of LLM variance against a human benchmark across sampling and persona conditions.
  2. [§5, Table 1, Appendix C] The systematic review is a central contribution (Contribution 3), and its headline counts (14/21 ground truth; 9/14 variance checked) depend on judgment calls that the authors acknowledge in Appendix C.1. No inter-coder reliability, coding sheet, or raw extracted data are provided, so a different reviewer could code borderline cases differently. For example, the assignment of heterogeneity requirements (High/Medium/Low) and the classification of 'Mean' and 'Var.' results (Aligned/Deviated/Mixed/Low) are not tied to explicit thresholds or verbatim excerpts. I ask the authors to make the coding transparent: include per-paper excerpts or a detailed coding table with the rationale for each dimension, and report inter-coder agreement if more than one coder was involved. Without this, the empirical support for the 'validation gap' claim is not independently verifiable, even though the conceptual argument may still stand.
  3. [§4.1; §5.3, Recommendation 2] The mean-variance framework is intuitive but under-specified. 'Variance' is defined only as 'the diversity and spread of behaviors' (§4.1), and no concrete metric is proposed for measuring it or for comparing against human baselines. This matters because Recommendation 2 asks researchers to 'report variance explicitly' and Recommendation 1 asks them to match validation depth to heterogeneity requirements; without an operationalization, these are hard to implement consistently across studies. I suggest the authors specify at least one example metric (e.g., variance or entropy of a target action distribution, inter-agent behavioral distance) and state how to determine whether variance is 'sufficient' relative to a human reference distribution, or explicitly delegate this to future work with a concrete proposal.
minor comments (4)
  1. [Abstract; §4.1] There are typographical issues: 'argues thatLLM-based' appears in the abstract and in §4.1, and 'Reynolds (1987)' should be 'Reynolds (1987)'s.'
  2. [Table 1; Appendix C.3] The Table 1 legend defines Aligned/Deviated/Mixed for Mean/Variance results but does not explain the entry 'Low' in the Var. column; Appendix C.3 defines 'Lower' for variance comparisons against ground truth, so the table's 'Low' should be either 'Lower' or explicitly defined.
  3. [Abstract; §5.1] The abstract's 'fewer than half explicitly assess behavioral variance' refers to the full sample (9/21), while §5.1 reports 'fewer explicitly assessed variance (9 of 14)' among ground-truth papers; consider clarifying the denominator in both places to avoid misreading.
  4. [References] The reference list contains duplicate entries for Hua et al. (2023/2024) and Fontana et al. (2024/2025); consider merging or cross-referencing to avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the boundary argument is conceptual and empirically grounded; self-citations are illustrative, not load-bearing.

full rationale

This is a position paper whose central claim—that LLM-based social simulations require clear boundaries and that LLMs' 'average persona' tendency limits behavioral diversity—is argued conceptually and supported by a systematic review, not derived by fitting parameters or by defining the conclusion into the premises. The average-persona phenomenon is supported by multiple independent external studies (e.g., Bisbee et al. 2024, Argyle et al. 2023, Xie et al. 2024, del Rio-Chanona et al. 2025), so the authors' citation of their own prior works (Wu et al. 2023, 2024) as illustrative examples is not load-bearing. The review's coding criteria involve acknowledged judgment calls, but that is a reliability limitation, not circularity: the counts do not reduce by construction to the paper's conclusions. The claim that temperature, sampling, or persona richness might restore variance is a substantive challenge to the strength of the 'fundamental limits' assertion, but it is a correctness concern, not a circular-derivation concern. No equation or fitted value is renamed as a prediction, and no self-cited uniqueness theorem forces the authors' choice. The paper is self-contained as an argumentative and empirical review, so no circularity is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new entities or fitted parameters. Its load-bearing premises are domain assumptions about the purpose of social simulation and the role of heterogeneity, plus the coding assumptions of the review.

assumptions (3)
  • domain assumption The primary objective of social simulation is explaining social patterns and generating hypotheses, not replication or prediction.
    Introduced in Section 2.1. The whole boundary argument depends on this normative stance; readers who think prediction is a valid primary goal will weight the conclusions differently.
  • domain assumption Research questions whose answers depend on distributional properties require higher agent heterogeneity.
    Stated in Appendix C.1 and used to assign High, Medium, Low heterogeneity requirements in Table 1. This is a reasonable domain assumption but not proven by the paper.
  • domain assumption Human population variance is the appropriate ground truth for judging LLM behavioral variance.
    Assumed in Section 3.3 and throughout the review. If the correct benchmark is not the full human distribution but some task-specific distribution, the low-variance verdict could shift.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LLM-Based Social Simulations Require a Boundary." pith.science (2026). https://pith.science/paper/RFPCXDWY

@misc{pith2026250619806,
  author       = {Pith},
  title        = {Pith review of: LLM-Based Social Simulations Require a Boundary},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RFPCXDWY}},
  note         = {Machine review of arXiv:2506.19806}
}
read the original abstract

This position paper argues that LLM-based social simulations require clear boundaries to make meaningful contributions to social science. While Large Language Models (LLMs) offer promising capabilities for simulating human behavior, their tendency to produce homogeneous outputs, acting as an "average persona", fundamentally limits their ability to capture the behavioral diversity essential for complex social dynamics. We examine why heterogeneity matters for social simulations and how current LLMs fall short, analyzing the relationship between mean alignment and variance in LLM-generated behaviors. Through a systematic review of representative studies, we find that validation practices often fail to match the heterogeneity requirements of research questions: while most papers include ground truth comparisons, fewer than half explicitly assess behavioral variance, and most that do report lower variance than human populations. We propose that researchers should: (1) match validation depth to the heterogeneity demands of their research questions, (2) explicitly report variance alongside mean alignment, and (3) constrain claims to collective-level qualitative patterns when variance is insufficient. Rather than dismissing LLM-based simulation, we advocate for a boundary-aware approach that ensures these methods contribute genuine insights to social science.

Figures

Figures reproduced from arXiv: 2506.19806 by the authors.

Figure 1
Figure 1. Overview of our claims. We value the goal of social simulations as a means to advance social science, e.g. by explaining social patterns, instead of focusing on “perfect” replication of real-world societies. We further examine possible simulation scenarios (e.g., aligned or misaligned means and variances) and advocate for a stronger emphasis on qualitative analysis of collective patterns. their tendency to act as an… view at source ↗
Figure 2
Figure 2. Distribution of chosen numbers by GPT-4 (blue) vs. humans (red), adapted from KBC (Wu et al., 2024). The LLM reproduces peak values (33, 50, 66) aligning with human choices, indicating aligned mean. However, the frequency of non-peak values is markedly lower than humans, highlighting low variance. human experiments or surveys) provides a basis for assessing whether LLM-generated behaviors exhibit realistic diversity… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Will Scaling Improve Social Simulation with LLMs?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.

  2. Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Grounding LLM agents with ACS demographics and ATUS time-use routines substantially improved reproduction of observed activity profiles under normal and heatwave conditions in Philadelphia, though over half of the hea...

  3. Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community

    cs.MA 2026-02 reject novelty 5.0 of 10

    Clustering of Moltbook submolt descriptions shows agent-created communities organize into human-mimetic, silicon-centric, and proto-economic themes, but the categories were partly prescribed by the analysis prompt.

Reference graph

Works this paper leans on

65 extracted references · 25 canonical work pages · cited by 3 Pith papers

  1. [1]

    Research Question Type and Heterogeneity Requirement Our classification draws on established frameworks in computational social science and ABM literature

    C.1. Research Question Type and Heterogeneity Requirement Our classification draws on established frameworks in computational social science and ABM literature. Following Squazzoni et al. (2014) and Edmonds et al. (2019), we recognize that different modeling purposes impose different requirements on agent heterogeneity. We adapt this idea to categorize re...

  2. [3]

    Mapping and influencing the po- litical ideology of large language models using synthetic personas.arXiv preprint arXiv:2412.14843,

    Bernardelle, P., Fr¨ohling, L., Civelli, S., Lunardi, R., Roitero, K., and Demartini, G. Mapping and influencing the po- litical ideology of large language models using synthetic personas.arXiv preprint arXiv:2412.14843,

  3. [6]

    T., Wu, J., and Liu, Z

    Chen, R., Li, Y ., Yang, J., Feng, Y ., Zhou, J. T., Wu, J., and Liu, Z. Identifying and mitigating social bias knowledge in language models. InFindings of the Association for Com- putational Linguistics: NAACL 2025, pp. 651–672,

  4. [7]

    Compost: Charac- terizing and evaluating caricature in llm simulations

    Cheng, M., Piccardi, T., and Yang, D. Compost: Charac- terizing and evaluating caricature in llm simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 10853–10875,

  5. [9]

    Simulating opinion dynamics with networks of llm-based agents

    Chuang, Y .-S., Goyal, A., Harlalka, N., Suresh, S., Hawkins, R., Yang, S., Shah, D., Hu, J., and Rogers, T. Simulating opinion dynamics with networks of llm-based agents. In Findings of the Association for Computational Linguistics: NAACL 2024, pp. 3326–3346, 2024a. Chuang, Y .-S., Nirunwiroj, K., Studdiford, Z., Goyal, A., Frigo, V ., Yang, S., Shah, D....

  6. [11]

    Behavioral and Topological Heterogeneities in Network Versions of Schelling's Segregation Model

    Deter, W. and Sayama, H. Behavioral and topological hetero- geneities in network versions of schelling’s segregation model.arXiv preprint arXiv:2408.05623,

  7. [12]

    Fontana, N., Pierri, F., and Aiello, L. M. Nicer than humans: How do large language models behave in the prisoner’s dilemma?arXiv preprint arXiv:2406.13605,

  8. [16]

    11 LLM-Based Social Simulations Require a Boundary Gandhi, K., Sadigh, D., and Goodman, N. D. Strategic reasoning with language models.arXiv preprint arXiv:2305.19165,

Show all 65 references
  1. [17]

    S3: Social-network simulation system with large language model-empowered agents.arXiv preprint arXiv:2307.14984,

    Gao, C., Lan, X., Lu, Z., Mao, J., Piao, J., Wang, H., Jin, D., and Li, Y . S3: Social-network simulation system with large language model-empowered agents.arXiv preprint arXiv:2307.14984,

  2. [18]

    Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,

    Ge, T., Chan, X., Wang, X., Yu, D., Mi, H., and Yu, D. Scaling synthetic data creation with 1,000,000,000 personas.arXiv preprint arXiv:2406.20094,

  3. [19]

    and Toubia, O

    Gui, G. and Toubia, O. The challenge of using llms to simulate human behavior: A causal inference perspective. arXiv preprint arXiv:2312.15524,

  4. [20]

    Can large language models play games? a case study of a self-play approach

    Guo, H., Liu, Z., Zhang, Y ., and Wang, Z. Can large language models play games? a case study of a self-play approach. arXiv preprint arXiv:2403.05632,

  5. [21]

    ”guinea pig trials” utilizing gpt: A novel smart agent-based modeling approach for studying firm competition and collusion.arXiv preprint arXiv:2308.10974,

    Han, X., Wu, Z., and Xiao, C. ”guinea pig trials” utilizing gpt: A novel smart agent-based modeling approach for studying firm competition and collusion.arXiv preprint arXiv:2308.10974,

  6. [22]

    M., Antunes, L., Pav´on Mestras, J., et al

    Hassan, S., Arroyo, J., Gal ´an Ordax, J. M., Antunes, L., Pav´on Mestras, J., et al. Asking the oracle: Introducing forecasting principles into agent-based modelling. Journal of artificial societies and social simulation. 2013, V . 16, n. 3,

  7. [23]

    J., and Xie, X

    Hu, Z., Lian, J., Xiao, Z., Xiong, M., Lei, Y ., Wang, T., Ding, K., Xiao, Z., Yuan, N. J., and Xie, X. Population-aligned persona generation for llm-based social simulation.arXiv preprint arXiv:2509.10127,

  8. [25]

    War and peace (waragent): Llm-based multi-agent simulation of world wars.arXiv preprint arXiv:2311.17227,

    Hua, W., Fan, L., Li, L., Mei, K., Ge, Y ., Hemphill, L., Zhang, Y ., et al. War and peace (waragent): Llm-based multi-agent simulation of world wars.arXiv preprint arXiv:2311.17227,

  9. [26]

    Social science meets llms: How reliable are large language models in so- cial simulations?arXiv preprint arXiv:2410.23426,

    Huang, Y ., Yuan, Z., Zhou, Y ., Guo, K., Wang, X., Zhuang, H., Sun, W., Sun, L., Wang, J., Ye, Y ., et al. Social science meets llms: How reliable are large language models in so- cial simulations?arXiv preprint arXiv:2410.23426,

  10. [27]

    Understanding llm agent behaviours via game theory: Strategy recognition, biases and multi-agent dynamics.arXiv preprint arXiv:2512.07462,

    Huynh, T.-K., Dao-Sy, D.-M., Cao, T.-B., Le, P.-H., Nguyen, H.-D., Nguyen-Lam, P.-Q., Nguyen-V o, M.-L., Pham, H.-P., Pham, P.-H., Than, T.-K., et al. Understanding llm agent behaviours via game theory: Strategy recognition, biases and multi-agent dynamics.arXiv preprint arXiv...

  11. [29]

    and T ¨ornberg, P

    Larooij, M. and T ¨ornberg, P. Do large language models solve the problems of agent-based modeling? a critical review of generative social simulations.arXiv preprint arXiv:2504.03274,

  12. [30]

    Llm generated persona is a promise with a catch.arXiv preprint arXiv:2503.16527, 2025a

    Li, A., Chen, H., Namkoong, H., and Peng, T. Llm generated persona is a promise with a catch.arXiv preprint arXiv:2503.16527, 2025a. Li, C. J., Wu, J., Mo, Z., Qu, A., Tang, Y ., Zhao, K. I., Gan, Y ., Fan, J., Yu, J., Zhao, J., et al. Simulating society requires simulating th...

  13. [31]

    Diversedia- logue: A methodology for designing chatbots with human- like diversity.arXiv preprint arXiv:2409.00262,

    Lin, X., Yu, X., Aich, A., Giorgi, S., and Ungar, L. Diversedia- logue: A methodology for designing chatbots with human- like diversity.arXiv preprint arXiv:2409.00262,

  14. [32]

    Evaluating large language model biases in persona-steered generation

    Liu, A., Diab, M., and Fried, D. Evaluating large language model biases in persona-steered generation. InFindings of the Association for Computational Linguistics ACL 2024, pp. 9832–9850, 2024a. Liu, X., Yu, H., Zhang, H., Xu, Y ., Lei, X., Lai, H., Gu, Y ., Ding, H., Men, K.,...

  15. [33]

    How does the heterogeneity of members affect the evolution of group opinions?Discrete Dynamics in Nature and Society, 2021(1):8827048,

    Lu, A., Ling, H., and Ding, Z. How does the heterogeneity of members affect the evolution of group opinions?Discrete Dynamics in Nature and Society, 2021(1):8827048,

  16. [35]

    Computational experiments meet large language model based agents: A survey and perspective.arXiv preprint arXiv:2402.00262,

    Ma, Q., Xue, X., Zhou, D., Yu, X., Liu, D., Zhang, X., Zhao, Z., Shen, Y ., Ji, P., Li, J., et al. Computational experiments meet large language model based agents: A survey and perspective.arXiv preprint arXiv:2402.00262,

  17. [36]

    J., and Ohno-Machado, L

    Ma, X., Zhu, R., Wang, Z., Xiong, J., Chen, Q., Tang, H., Camp, L. J., and Ohno-Machado, L. Enhancing patient- centric communication: Leveraging llms to simulate patient perspectives.arXiv preprint arXiv:2501.06964,

  18. [37]

    Towards a holistic landscape of situated theory of mind in large language models.arXiv preprint arXiv:2310.19619,

    Ma, Z., Sansom, J., Peng, R., and Chai, J. Towards a holistic landscape of situated theory of mind in large language models.arXiv preprint arXiv:2310.19619,

  19. [40]

    and Xu, W

    Naous, T. and Xu, W. On the origin of cultural biases in language models: From pre-training data to linguistic phenomena.arXiv preprint arXiv:2501.04662,

  20. [41]

    Beyond demographics: Fine-tuning large language models to predict individuals’ subjective text perceptions.arXiv preprint arXiv:2502.20897,

    Orlikowski, M., Pei, J., R¨ottger, P., Cimiano, P., Jurgens, D., and Hovy, D. Beyond demographics: Fine-tuning large language models to predict individuals’ subjective text perceptions.arXiv preprint arXiv:2502.20897,

  21. [42]

    S., Zou, C

    Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Willer, R., Liang, P., and Bernstein, M. S. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109,

  22. [45]

    P., Galley, M., Caruana, R., and Gao, J

    Singh, C., Inala, J. P., Galley, M., Caruana, R., and Gao, J. Rethinking interpretability in the era of large language models.arXiv preprint arXiv:2402.01761,

  23. [46]

    R., Jain, R., and Mehta, S

    Surve, A., Rathod, A., Surana, M., Malpani, G., Shamraj, A., Sankepally, S. R., Jain, R., and Mehta, S. S. Mul- tiagent simulators for social networks.arXiv preprint arXiv:2311.14712,

  24. [47]

    Gensim: A general social simulation platform with large language model based agents

    Tang, J., Gao, H., Pan, X., Wang, L., Tan, H., Gao, D., Chen, Y ., Chen, X., Lin, Y ., Li, Y ., et al. Gensim: A general social simulation platform with large language model based agents. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ...

  25. [48]

    Systematic biases in llm simulations of debates

    Taubenfeld, A., Dover, Y ., Reichart, R., and Goldstein, A. Systematic biases in llm simulations of debates. InPro- ceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 251–267,

  26. [50]

    Large language models as urban residents: An llm agent framework for personal mobility generation.Advances in Neural Information Processing Systems, 37:124547–124574, 2024a

    Wang, J., Jiang, R., Yang, C., Wu, Z., Onizuka, M., Shibasaki, R., Koshizuka, N., Xiao, C., et al. Large language models as urban residents: An llm agent framework for personal mobility generation.Advances in Neural Information Processing Systems, 37:124547–124574, 2024a. 15 L...

  27. [51]

    From chatgpt to deepseek: Can llms simulate humanity?arXiv preprint arXiv:2502.18210, 2025c

    Wang, Q., Tang, Z., and He, B. From chatgpt to deepseek: Can llms simulate humanity?arXiv preprint arXiv:2502.18210, 2025c. Wang, Q., Wu, J., Tang, Z., Luo, B., Chen, N., Chen, W., and He, B. What limits llm-based human simulation: Llms or our design?arXiv preprint arXiv:2501....

  28. [52]

    Smart agent-based modeling: On the use of large language models in computer simulations.arXiv preprint arXiv:2311.06330,

    Wu, Z., Peng, R., Han, X., Zheng, S., Zhang, Y ., and Xiao, C. Smart agent-based modeling: On the use of large language models in computer simulations.arXiv preprint arXiv:2311.06330,

  29. [53]

    Shall we team up: Exploring spontaneous cooperation of competing llm agents

    Wu, Z., Peng, R., Zheng, S., Liu, Q., Han, X., Kwon, B., Onizuka, M., Tang, S., and Xiao, C. Shall we team up: Exploring spontaneous cooperation of competing llm agents. InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 5163–5186,

  30. [55]

    Evaluating and enhancing llms agent based on theory of mind in guandan: A multi-player cooperative game under imperfect information.arXiv preprint arXiv:2408.02559,

    Yim, Y ., Chan, C., Shi, T., Deng, Z., Fan, W., Zheng, T., and Song, Y . Evaluating and enhancing llms agent based on theory of mind in guandan: A multi-player cooperative game under imperfect information.arXiv preprint arXiv:2408.02559,

  31. [56]

    Exploring collaboration mechanisms for llm agents: A social psychology view.arXiv preprint arXiv:2310.02124,

    Zhang, J., Xu, X., Zhang, N., Liu, R., Hooi, B., and Deng, S. Exploring collaboration mechanisms for llm agents: A social psychology view.arXiv preprint arXiv:2310.02124,

  32. [57]

    Don’t go to extremes: Revealing the excessive sensitivity and calibration limitations of llms in implicit hate speech detection

    Zhang, M., He, J., Ji, T., and Lu, C.-T. Don’t go to extremes: Revealing the excessive sensitivity and calibration limitations of llms in implicit hate speech detection. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  33. [58]

    K-level reasoning with large language models

    Zhang, Y ., Mao, S., Ge, T., Wang, X., Xia, Y ., Lan, M., and Wei, F. K-level reasoning with large language models. arXiv e-prints, pp. arXiv–2402, 2024b. Zhang, Z., Hu, F., Lee, J., Shi, F., Kordjamshidi, P., Chai, J., and Ma, Z. Do vision-language models represent space and ...

  34. [59]

    H., Wang, X., Zhu, H., Wang, W., and Sap, M

    Zhou, J., Huang, J.-t., Zhou, X., Lam, M. H., Wang, X., Zhu, H., Wang, W., and Sap, M. The pimmur principles: Ensuring validity in collective behavior of llm societies. arXiv preprint arXiv:2509.18052,

  35. [60]

    Sotopia: Interactive evaluation for social intelligence in language agents.arXiv preprint arXiv:2310.11667,

    Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y ., Fried, D., Neubig, G., et al. Sotopia: Interactive evaluation for social intelligence in language agents.arXiv preprint arXiv:2310.11667,

  36. [61]

    apophenia

    17 LLM-Based Social Simulations Require a Boundary A. Related Works A.1. Computational Social Science Social phenomena typically arise from the interactions of intelligent, adaptive agents under dynamic conditions (Eidelson, 1997; San Miguel et al., 2012). Even when we fully u...

  37. [63]

    (3) Exploratory studies have demonstrated human-like behavior, with performance approaching that of humans in certain experiments (Anthis et al., 2025)

    and improving the fidelity of complex behaviors such as interaction, collaboration, and gaming (Ma et al., 2024). (3) Exploratory studies have demonstrated human-like behavior, with performance approaching that of humans in certain experiments (Anthis et al., 2025). However, r...

  38. [65]

    Would the research question still be answerable if all agents behaved identically at the population mean?

    • Low: Questions primarily concerning equilibrium existence or central tendency dynamics, where the specific distribution shape matters less. 19 LLM-Based Social Simulations Require a Boundary •Medium: Questions involving some distributional aspects but where central tendencie...

  39. [100]

    Can markets clear?

    C.3. Alignment Assessment Mean Alignmentwas marked as checked (✓) if the paper explicitly compared average LLM agent behavior against human baselines. Results were categorized as: •Aligned: LLM mean behavior is statistically or qualitatively consistent with human average. •Dev...

  40. [1971]

    Personality traits in large language models.arXiv preprint arXiv:2307.00184,

    Serapio-Garc´ıa, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., and Matari´c, M. Personality traits in large language models.arXiv preprint arXiv:2307.00184,

  41. [1988]

    United in diversity? contextual biases in llm-based predictions of the 2024 european parliament elections.arXiv preprint arXiv:2409.09045,

    von der Heyde, L., Haensch, A.-C., and Wenz, A. United in diversity? contextual biases in llm-based predictions of the 2024 european parliament elections.arXiv preprint arXiv:2409.09045,

  42. [1996]

    K., Bhatia, S., and Chakraborty, T

    Chatterjee, A., Renduchintala, H. K., Bhatia, S., and Chakraborty, T. Posix: A prompt sensitivity index for large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 14550–14565,

  43. [1999]

    and Schelling’s segregation model (Schelling, 1971), which illustrate how wealth gaps or segregation patterns can emerge from individual interactions. Despite disagreements and inconsistencies within social science theories, many works agree that social interaction is the fund...

  44. [2005]

    SIS 2005., pp. 201–208. IEEE,

  45. [2006]

    Modeling and math- ematical analysis of swarms of microscopic robots

    Galstyan, A., Hogg, T., and Lerman, K. Modeling and math- ematical analysis of swarms of microscopic robots. In Proceedings 2005 IEEE Swarm Intelligence Symposium,

  46. [2009]

    Explaining large language models decisions using shapley values.arXiv preprint arXiv:2404.01332,

    Mohammadi, B. Explaining large language models decisions using shapley values.arXiv preprint arXiv:2404.01332,

  47. [2010]

    Emergence of social norms in generative agent societies: principles and architecture.arXiv preprint arXiv:2403.08251,

    14 LLM-Based Social Simulations Require a Boundary Ren, S., Cui, Z., Song, R., Wang, Z., and Hu, S. Emergence of social norms in generative agent societies: principles and architecture.arXiv preprint arXiv:2403.08251,

  48. [2011]

    Specializing large language models to simulate survey response distributions for global populations.arXiv preprint arXiv:2502.07068,

    10 LLM-Based Social Simulations Require a Boundary Cao, Y ., Liu, H., Arora, A., Augenstein, I., R¨ottger, P., and Hershcovich, D. Specializing large language models to simulate survey response distributions for global populations.arXiv preprint arXiv:2502.07068,

  49. [2014]

    Are large language models (llms) good social predictors? InFindings of the Association for Computational Linguistics: EMNLP 2024, pp

    Yang, K., Li, H., Wen, H., Peng, T.-Q., Tang, J., and Liu, H. Are large language models (llms) good social predictors? InFindings of the Association for Computational Linguistics: EMNLP 2024, pp. 2718–2730, 2024a. Yang, Y ., Duan, H., Liu, J., and Tam, K. Y . Llm-measure: Gene...

  50. [2015]

    M., Pangallo, M., and Hommes, C

    del Rio-Chanona, R. M., Pangallo, M., and Hommes, C. Can generative ai agents behave like humans? evidence from laboratory market experiments.arXiv preprint arXiv:2505.07457,

  51. [2017]

    G., Ortu, F., Strausz, A., Sachan, M., Mihalcea, R., et al

    Jin, Z., Kleiman-Weiner, M., Piatti, G., Levine, S., Liu, J., Adauto, F. G., Ortu, F., Strausz, A., Sachan, M., Mihalcea, R., et al. Multilingual trolley problems for language models. InPluralistic Alignment Workshop at NeurIPS 2024,

  52. [2018]

    R., Liu, R., Richardson, S

    Anthis, J. R., Liu, R., Richardson, S. M., Kozlowski, A. C., Koch, B., Evans, J., Brynjolfsson, E., and Bernstein, M. Llm social simulations are a promising research method. arXiv preprint arXiv:2504.02234,

  53. [2021]

    Lu, S. E. Strategic interactions between large language models-based agents in beauty contests.arXiv preprint arXiv:2404.08492,

  54. [2022]

    From individual to society: A survey on social simulation driven by large language model-based agents.arXiv preprint arXiv:2412.03563,

    Mou, X., Ding, X., He, Q., Wang, L., Liang, J., Zhang, X., Sun, L., Lin, J., Zhou, J., Huang, X., et al. From individual to society: A survey on social simulation driven by large language model-based agents.arXiv preprint arXiv:2412.03563,

  55. [2023]

    Observing micromotives and macrobehavior of large language models.arXiv preprint arXiv:2412.10428,

    Cheng, Y ., Qu, X., Goldsack, T., Lin, C., and Chen, C.-C. Observing micromotives and macrobehavior of large language models.arXiv preprint arXiv:2412.10428,

  56. [2024]

    Can we count on llms? the fixed-effect fallacy and claims of gpt-4 capabilities.arXiv preprint arXiv:2409.07638,

    Ball, T., Chen, S., and Herley, C. Can we count on llms? the fixed-effect fallacy and claims of gpt-4 capabilities.arXiv preprint arXiv:2409.07638,

  57. [2025]

    Anomalous fluctua- tions in minority games and related multi-agent models of financial markets.arXiv preprint physics/0608091,

    Galla, T., Mosetti, G., and Zhang, Y .-C. Anomalous fluctua- tions in minority games and related multi-agent models of financial markets.arXiv preprint physics/0608091,

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.