REVIEW 16 cited by
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Can we avoid wars at the crossroads of history? This question has been pursued by individuals, scholars, policymakers, and organizations throughout human history. In this research, we attempt to answer the question based on the recent advances of Artificial Intelligence (AI) and Large Language Models (LLMs). We propose \textbf{WarAgent}, an LLM-powered multi-agent AI system, to simulate the participating countries, their decisions, and the consequences, in historical international conflicts, including the World War I (WWI), the World War II (WWII), and the Warring States Period (WSP) in Ancient China. By evaluating the simulation effectiveness, we examine the advancements and limitations of cutting-edge AI systems' abilities in studying complex collective human behaviors such as international conflicts under diverse settings. In these simulations, the emergent interactions among agents also offer a novel perspective for examining the triggers and conditions that lead to war. Our findings offer data-driven and AI-augmented insights that can redefine how we approach conflict resolution and peacekeeping strategies. The implications stretch beyond historical analysis, offering a blueprint for using AI to understand human history and possibly prevent future international conflicts. Code and data are available at \url{https://github.com/agiresearch/WarAgent}.
Forward citations
Cited by 16 Pith papers
-
On Path to Multimodal Historical Reasoning: HistBench and HistAgent
HistAgent, a history-specialized agent, scores 27.54% pass@1 and 36.47% pass@2 on the new 414-question HistBench benchmark, surpassing generalist agents tested on the same data.
-
Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems
Explicit positive or negative relationship cues in multi-agent prompts act mainly as convergence pressure, increasing agreement without reliably improving answer correctness.
-
No One Wins in Nuclear War: A Social Simulation of Military Decision-making
WOPR is a deterministic, replay-checkable rules engine that turns the card game Nuclear War into a social-simulation testbed for military decision-making.
-
Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
Eco3S packages co-evolving environments, checkpoint counterfactuals, and auto-refinement into one LLM agent-based simulation platform, demonstrated on canal-rebellion, state-formation, and information-spread cases.
-
Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling
LLM agents with demographic profiles reproduce CPT-style risk attitudes in route choice and yield fitted parameters (α=0.4, β=0.64, λ=1.43) that predict human data competitively.
-
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges
Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.
-
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.
-
Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
Co-Saving cuts token usage by roughly half in multi-agent software development by injecting learned shortcut instructions that bypass intermediate reasoning steps, while slightly improving a composite code-quality score.
-
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
The DynToM benchmark shows ten LLMs average 33.0% accuracy versus 77.7% for humans, and models lose the most accuracy on questions about mental-state changes across scenarios.
-
Ethical Considerations of Large Language Models in Game Playing
In Werewolf games, LLM agents change their kills, votes, and trust scores based on explicit gender labels and even based on gender-implied first names, behaving differently for male and female players.
-
Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences
LLM agents can run a simulated decision conference, and a dedicated agreement-detection agent helps the debate cover topics that match a real expert workshop.
-
Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design
A heuristic framework that decomposes known algorithms into typed LLM-agent subtasks lifts small-model accuracy on knapsack and assignment problems from near-zero to high levels after fixing one bottleneck agent.
-
Artificial Intelligence and Civil Discourse: How LLMs Moderate Climate Change Conversations
LLM replies to climate change posts are more emotionally neutral and lower in intensity than the human posts they respond to.
-
AI Agent Behavioral Science
AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.
-
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
ACBench tests compressed LLMs on agentic tasks and finds 4-bit quantization keeps tool use and workflow generation strong while hurting real-world application performance.
-
Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods
LLMs extend, rather than replace, classical social science methods, with a proposed three-tier bias framework for LLM-augmented surveys.
Discussion (0). Continue with ORCID to comment.