Pith. sign in

REVIEW 3 cited by

Emergent social conventions and collective bias in LLM populations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08948 v2 pith:N4CEQDTJ submitted 2024-10-11 cs.MA cs.AIcs.CYphysics.soc-ph

classification cs.MAcs.AIcs.CYphysics.soc-ph
keywords socialconventionsagentspopulationsbiascollectivelanguageresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Social conventions are the backbone of social coordination, shaping how individuals form a group. As growing populations of artificial intelligence (AI) agents communicate through natural language, a fundamental question is whether they can bootstrap the foundations of a society. Here, we present experimental results that demonstrate the spontaneous emergence of universally adopted social conventions in decentralized populations of large language model (LLM) agents. We then show how strong collective biases can emerge during this process, even when agents exhibit no bias individually. Last, we examine how committed minority groups of adversarial LLM agents can drive social change by imposing alternative social conventions on the larger population. Our results show that AI systems can autonomously develop social conventions without explicit programming and have implications for designing AI systems that align, and remain aligned, with human values and societal goals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

    cs.CL 2026-07 accept novelty 7.0 of 10

    Across six open-weight LLMs and seven datasets, a speaker-free wrong-answer assertion alone flips 66.5% of initially correct answers, versus 10.3% for a plain re-ask; source labels mainly add a modest increment above ...

  2. How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-4 impersonating Reddit users in 2016 election threads produces comments that lean toward consensus and are semantically separable from real human comments.

  3. Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability

    cs.AI 2025-09 conditional novelty 3.0 of 10

    An LLM output can be verified by regenerating a few randomly chosen segments under identical hardware, with a tunable detection probability and 12.4x speedup over full regeneration.

Pith tools