Pith. sign in

REVIEW 8 cited by

Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.10264 v5 pith:2ZKJYJY7 submitted 2022-08-18 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagemodelshumansimulatingbehaviordifferentexperimentfindings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a language model's simulation of a specific human behavior. Unlike the Turing Test, which involves simulating a single arbitrary individual, a TE requires simulating a representative sample of participants in human subject research. We carry out TEs that attempt to replicate well-established findings from prior studies. We design a methodology for simulating TEs and illustrate its use to compare how well different language models are able to reproduce classic economic, psycholinguistic, and social psychology experiments: Ultimatum Game, Garden Path Sentences, Milgram Shock Experiment, and Wisdom of Crowds. In the first three TEs, the existing findings were replicated using recent models, while the last TE reveals a "hyper-accuracy distortion" present in some language models (including ChatGPT and GPT-4), which could affect downstream applications in education and the arts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 124 citations worldwide. Full citation record

  1. Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment

    cs.CL 2026-06 unverdicted novelty 6.5 of 10

    On UPBench, 25 LLMs show a non-monotonic planning curve—strong Remember/Analyze, weak Understand/Evaluate—with four failure modes that support differential, not blanket, AI delegation.

  2. Large language models replicate and predict human cooperation across experiments in game theory

    cs.AI 2025-11 conditional novelty 6.0 of 10

    Llama-3.1-8B with a multi-step reasoning-and-filter prompt reproduces human cooperation rates across 121 dyadic games (MSD=0.031, r=0.89), outperforming Nash-equilibrium predictions (MSD=0.096, r=0.78).

  3. Super-additive Cooperation in Language Model Agents

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Language model agents cooperate more in a prisoner's dilemma when repeated interactions and inter-team competition are combined, but only for some models.

  4. The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games

    cs.AI 2025-06 conditional novelty 6.0 of 10

    In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...

  5. Modeling Earth-Scale Human-Like Societies with One Billion Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.

  6. Using AI to replicate human experimental results: a motion study

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Across four motion-verb psycholinguistic tasks, ChatGPT o1 responses correlated strongly with human judgments (Spearman rho around .69 to .96), but methodological gaps limit the strength of the replication claim.

  7. Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods

    cs.AI 2025-05 accept novelty 4.0 of 10

    LLMs extend, rather than replace, classical social science methods, with a proposed three-tier bias framework for LLM-augmented surveys.

  8. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    cs.LG 2025-02

Pith tools