Pith. sign in

REVIEW 12 cited by

Strategic Reasoning with Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.19165 v1 pith:3X3EI2GP submitted 2023-05-30 cs.AI cs.CLcs.GTcs.HC

classification cs.AIcs.CLcs.GTcs.HC
keywords strategicreasoningagentsapproachgameslanguagellmsscenarios
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Strategic reasoning enables agents to cooperate, communicate, and compete with other agents in diverse situations. Existing approaches to solving strategic games rely on extensive training, yielding strategies that do not generalize to new scenarios or games without retraining. Large Language Models (LLMs), with their ability to comprehend and generate complex, context-rich language, could prove powerful as tools for strategic gameplay. This paper introduces an approach that uses pretrained LLMs with few-shot chain-of-thought examples to enable strategic reasoning for AI agents. Our approach uses systematically generated demonstrations of reasoning about states, values, and beliefs to prompt the model. Using extensive variations of simple matrix games, we show that strategies that are derived based on systematically generated prompts generalize almost perfectly to new game structures, alternate objectives, and hidden information. Additionally, we demonstrate our approach can lead to human-like negotiation strategies in realistic scenarios without any extra training or fine-tuning. Our results highlight the ability of LLMs, guided by systematic reasoning demonstrations, to adapt and excel in diverse strategic scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games

    cs.GT 2026-07 conditional novelty 6.0 of 10

    A two-feature game embedding (Nash entropy and best-response switching) predicts cross-game transfer of fine-tuned LLMs on held-out games, outperforming game identity and published structural embeddings.

  2. When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems

    cs.MA 2026-01 unverdicted novelty 6.0 of 10

    Role-based personas in multi-agent LLM systems suppress payoff-aligned behavior, shifting equilibrium selection by up to 90 percentage points in Tragedy of the Commons versus Green Transition scenarios even with full ...

  3. We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new LLM-generated dataset (PACT) of 8,687 personality-tagged, argumentation-annotated tourism negotiations, plus a three-part benchmark in which fine-tuned models beat zero-shot and human-human-data baselines.

  4. Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy

    cs.AI 2025-08 conditional novelty 6.0 of 10

    An evaluation harness lets off-the-shelf local LLMs, including a 24B model, play full-press Diplomacy without fine-tuning.

  5. Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic Reasoning

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A hierarchical framework of specialized imitation agents plus a strategic planner improves win rates and cuts LLM calls in text-based StarCraft II across all race matchups.

  6. Are Reasoning Models More Prone to Hallucination?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Post-training pipeline choice (SFT+RL vs RL-only vs SFT-only) reliably shifts hallucination rates in large reasoning models on fact-seeking benchmarks.

  7. Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A prompt-shared hierarchical offline RL pipeline for LLM agents improves long-horizon task scores on ScienceWorld and ALFWorld over non-hierarchical baselines.

  8. When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across morally framed prisoner's dilemmas and public goods games, none of nine LLMs consistently chooses the ethical action when it conflicts with payoff, with cooperation rates from 7.9% to 76.3%.

  9. InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    InstantEdit combines RectifiedFlow inversion, latent injection, disentangled prompt guidance, and Canny ControlNet to do fast few-step text-guided image editing with content preservation.

  10. Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making

    cs.AI 2025-06 conditional novelty 5.0 of 10

    LLMs reproduce human heuristics like switching after a loss and cooperating when future rounds loom, but apply them more rigidly and adapt less than humans.

  11. Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets

    cs.AI 2025-05 conditional novelty 5.0 of 10

    AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.

  12. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Reinforcing the two experts most correlated with thinking tokens improves reasoning accuracy and efficiency in MoE large reasoning models, with gains of up to 10 points on AIME benchmarks.

Pith tools