Pith. sign in

REVIEW 3 cited by

Simulating Human Strategic Behavior: Comparing Single and Multi-agent LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08189 v2 pith:UOL3AFM3 submitted 2024-02-13 cs.HC

classification cs.HC
keywords llmsreasoninghumansimulatepeoplestrategicmulti-agentsingle
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When creating policies, plans, or designs for people, it is challenging for designers to foresee all of the ways in which people may reason and behave. Recently, Large Language Models (LLMs) have been shown to be able to simulate human reasoning. We extend this work by measuring LLMs ability to simulate strategic reasoning in the ultimatum game, a classic economics bargaining experiment. Experimental evidence shows human strategic reasoning is complex; people will often choose to punish other players to enforce social norms even at personal expense. We test if LLMs can replicate this behavior in simulation, comparing two structures: single LLMs and multi-agent systems. We compare their abilities to (1) simulate human-like reasoning in the ultimatum game, (2) simulate two player personalities, greedy and fair, and (3) create robust strategies that are logically complete and consistent with personality. Our evaluation shows that multi-agent systems are more accurate than single LLMs (88 percent vs. 50 percent) in simulating human reasoning and actions for personality pairs. Thus, there is potential to use LLMs to simulate human strategic reasoning to help decision and policy-makers perform preliminary explorations of how people behave in systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large language models replicate and predict human cooperation across experiments in game theory

    cs.AI 2025-11 conditional novelty 6.0 of 10

    Llama-3.1-8B with a multi-step reasoning-and-filter prompt reproduces human cooperation rates across 121 dyadic games (MSD=0.031, r=0.89), outperforming Nash-equilibrium predictions (MSD=0.096, r=0.78).

  2. MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications

    cs.MA 2025-06 conditional novelty 5.0 of 10

    MATE is an open-source multi-agent system for accessibility that uses a fine-tuned BERT model, trained on a new AI-generated dataset, to recognize and execute modality conversion tasks.

  3. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

Pith tools