REVIEW 13 cited by
LLM Social Simulations Are a Promising Research Method
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Accurate and verifiable large language model (LLM) simulations of human research subjects promise an accessible data source for understanding human behavior and training new AI systems. However, results to date have been limited, and few social scientists have adopted this method. In this position paper, we argue that the promise of LLM social simulations can be achieved by addressing five tractable challenges. We ground our argument in a review of empirical comparisons between LLMs and human research subjects, commentaries on the topic, and related work. We identify promising directions, including context-rich prompting and fine-tuning with social science datasets. We believe that LLM social simulations can already be used for pilot and exploratory studies, and more widespread use may soon be possible with rapidly advancing LLM capabilities. Researchers should prioritize developing conceptual models and iterative evaluations to make the best use of new AI systems.
Forward citations
Cited by 13 Pith papers
-
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
LLM probability estimates violate the law of total probability across partitions, and subgroup-aggregated estimates often beat direct population-level estimates (the macro fallacy).
-
Will Scaling Improve Social Simulation with LLMs?
Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.
-
Tailored untruths: How personalisation challenges LLM safeguards
A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.
-
Informing AI Policy Assessment using Large-Scale Simulation of Interventions
A genetic algorithm optimizes weighted combinations of LLM-perceived harm mitigation, expert costs, and participatory scores over stakeholder-action pairs to surface viable AI policy packages for media harms.
-
The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions
LLMs attribute moral responsibility like humans but refuse to act on it in scarce-resource allocation, defaulting to random choice instead of favoring the less-culpable patient.
-
Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions
Language model agents with personality and theory-of-mind prompts reproduce human third-party punishment and gossip effects, and predict lower anonymous punishment and higher cooperation after group discussion.
-
Social Scientists on the Role of AI in Research
Randomized survey wording makes social scientists report more familiarity but less trust in "AI" than in "machine learning", with ethical concerns concentrated on generative AI.
-
Aligning LLM with human travel choices: a persona-based embedding learning approach
A persona-based embedding learning framework aligns LLM predictions with human travel mode choices, outperforming MNL and few-shot LLM baselines on the Swissmetro dataset.
-
Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour
AXIS couples an LLM with a multi-agent simulator to produce counterfactual action explanations, and reports higher judged correctness and goal-prediction accuracy than baselines on ten autonomous-driving scenarios.
-
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
Policy claims from LLM community simulations should be stated as probabilities of necessary and sufficient causation, mapped to stakeholder needs and conditioned on simulator fidelity.
-
Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community
Clustering of Moltbook submolt descriptions shows agent-created communities organize into human-mimetic, silicon-centric, and proto-economic themes, but the categories were partly prescribed by the analysis prompt.
-
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
A design proposal, not a tested system: a dual-view simulation loop in which an LLM-simulated user issues reflective instructions when about to disengage, and an instruction-following recommender adapts to maximize si...
-
Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents
In three simulation runs, LLM agents primed with a swine flu article reduced social activity compared with controls, but the effect is confounded by prompt instructions and no inferential statistics are reported.
Discussion (0). Continue with ORCID to comment.