Pith. sign in

REVIEW 13 cited by

LLM Social Simulations Are a Promising Research Method

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02234 v2 pith:LMQXNPCT submitted 2025-04-03 cs.HC cs.AIcs.CLcs.CY

classification cs.HCcs.AIcs.CLcs.CY
keywords socialsimulationshumanresearchmethodpromisepromisingsubjects
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurate and verifiable large language model (LLM) simulations of human research subjects promise an accessible data source for understanding human behavior and training new AI systems. However, results to date have been limited, and few social scientists have adopted this method. In this position paper, we argue that the promise of LLM social simulations can be achieved by addressing five tractable challenges. We ground our argument in a review of empirical comparisons between LLMs and human research subjects, commentaries on the topic, and related work. We identify promising directions, including context-rich prompting and fine-tuning with social science datasets. We believe that LLM social simulations can already be used for pilot and exploratory studies, and more widespread use may soon be possible with rapidly advancing LLM capabilities. Researchers should prioritize developing conceptual models and iterative evaluations to make the best use of new AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    LLM probability estimates violate the law of total probability across partitions, and subgroup-aggregated estimates often beat direct population-level estimates (the macro fallacy).

  2. Will Scaling Improve Social Simulation with LLMs?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.

  3. Tailored untruths: How personalisation challenges LLM safeguards

    cs.CL 2025-10 conditional novelty 7.0 of 10

    A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.

  4. Informing AI Policy Assessment using Large-Scale Simulation of Interventions

    cs.CY 2026-04 conditional novelty 6.5 of 10

    A genetic algorithm optimizes weighted combinations of LLM-perceived harm mitigation, expert costs, and participatory scores over stakeholder-action pairs to surface viable AI policy packages for media harms.

  5. The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

    cs.CY 2026-08 conditional novelty 6.0 of 10

    LLMs attribute moral responsibility like humans but refuse to act on it in scarce-resource allocation, defaulting to random choice instead of favoring the less-culpable patient.

  6. Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions

    cs.MA 2025-07 conditional novelty 6.0 of 10

    Language model agents with personality and theory-of-mind prompts reproduce human third-party punishment and gossip effects, and predict lower anonymous punishment and higher cooperation after group discussion.

  7. Social Scientists on the Role of AI in Research

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Randomized survey wording makes social scientists report more familiarity but less trust in "AI" than in "machine learning", with ethical concerns concentrated on generative AI.

  8. Aligning LLM with human travel choices: a persona-based embedding learning approach

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A persona-based embedding learning framework aligns LLM predictions with human travel mode choices, outperforming MNL and few-shot LLM baselines on the Swissmetro dataset.

  9. Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour

    cs.AI 2025-05 conditional novelty 6.0 of 10

    AXIS couples an LLM with a multi-agent simulator to produce counterfactual action explanations, and reports higher judged correctness and goal-prediction accuracy than baselines on ten autonomous-driving scenarios.

  10. From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities

    cs.CL 2026-04 conditional novelty 5.0 of 10

    Policy claims from LLM community simulations should be stated as probabilities of necessary and sufficient causation, mapped to stakeholder needs and conditioned on simulator fidelity.

  11. Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community

    cs.MA 2026-02 reject novelty 5.0 of 10

    Clustering of Moltbook submolt descriptions shows agent-created communities organize into human-mimetic, silicon-centric, and proto-economic themes, but the categories were partly prescribed by the analysis prompt.

  12. RecoWorld: Building Simulated Environments for Agentic Recommender Systems

    cs.IR 2025-09 conditional novelty 5.0 of 10

    A design proposal, not a tested system: a dual-view simulation loop in which an LLM-simulated user issues reflective instructions when about to disengage, and an instruction-following recommender adapts to maximize si...

  13. Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents

    physics.soc-ph 2025-06 reject novelty 5.0 of 10

    In three simulation runs, LLM agents primed with a swine flu article reduced social activity compared with controls, but the effect is confounded by prompt instructions and no inferential statistics are reported.

Pith tools