Pith. sign in

REVIEW 4 cited by

Cultural Evolution of Cooperation among LLM Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.10270 v1 pith:XJWJXME3 submitted 2024-12-13 cs.MA cs.AI

classification cs.MAcs.AI
keywords agentsacrossevolutionbehaviorclassclaudecooperationdeployment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) provide a compelling foundation for building generally-capable AI agents. These agents may soon be deployed at scale in the real world, representing the interests of individual humans (e.g., AI assistants) or groups of humans (e.g., AI-accelerated corporations). At present, relatively little is known about the dynamics of multiple LLM agents interacting over many generations of iterative deployment. In this paper, we examine whether a "society" of LLM agents can learn mutually beneficial social norms in the face of incentives to defect, a distinctive feature of human sociality that is arguably crucial to the success of civilization. In particular, we study the evolution of indirect reciprocity across generations of LLM agents playing a classic iterated Donor Game in which agents can observe the recent behavior of their peers. We find that the evolution of cooperation differs markedly across base models, with societies of Claude 3.5 Sonnet agents achieving significantly higher average scores than Gemini 1.5 Flash, which, in turn, outperforms GPT-4o. Further, Claude 3.5 Sonnet can make use of an additional mechanism for costly punishment to achieve yet higher scores, while Gemini 1.5 Flash and GPT-4o fail to do so. For each model class, we also observe variation in emergent behavior across random seeds, suggesting an understudied sensitive dependence on initial conditions. We suggest that our evaluation regime could inspire an inexpensive and informative new class of LLM benchmarks, focussed on the implications of LLM agent deployment for the cooperative infrastructure of society.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations

    cs.MA 2025-10 conditional novelty 6.0 of 10

    In networked LLM agents, simply informing influence-operation agents of their teammates' identities produces coordination nearly as strong as collective deliberation and voting.

  2. Modeling Earth-Scale Human-Like Societies with One Billion Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.

  3. Emergence of Reputation-Based Cooperation in LLM Agents

    cs.MA 2026-08 conditional novelty 5.0 of 10

    AI agents evolve donation strategies resembling Image Scoring, and the steepness of their discrimination against uncooperative opponents predicts resistance to free-riders.

  4. From Connectivity to Autonomy: The Dawn of Self-Evolving Communication Systems

    eess.SY 2025-05 conditional novelty 3.0 of 10

    The paper sketches a four-layer AI-enabled architecture for self-evolving 6G networks and a roadmap to implement it.

Pith tools