Pith. sign in

REVIEW 20 cited by

Evaluating Large Language Models in Theory of Mind Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.02083 v7 pith:3H47IKR7 submitted 2023-02-04 cs.CL cs.CYcs.HC

classification cs.CLcs.CYcs.HC
keywords taskslanguagemodelssolvedacrossbatteryconsideredfalse-belief
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The battery included 640 prompts spread across 40 diverse tasks, each one including a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. To solve a single task, a model needed to correctly answer 16 prompts across all eight scenarios. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. We explore the potential interpretation of these findings, including the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 176 citations worldwide. Full citation record

  1. Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.

  2. MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Credit assignment via LMM pairwise comparisons plus Bradley–Terry rank aggregation and potential-based shaping improves cooperative MARL under sparse rewards and dynamic agent counts.

  3. AgentSociety 2: An Integrated Research Environment for Executable Social Science

    cs.CY 2026-06 unverdicted novelty 6.0 of 10

    An integrated LLM-agent environment runs social-science experiments from hypothesis to manuscript, reproducing several known human patterns while failing on others (implicit self-bias, free-riding decay, norm collapse).

  4. Tacit Coordination of Large Language Models

    cs.GT 2026-01 conditional novelty 6.0 of 10

    Across 20+ open-source LLMs, tacit coordination in focal-point games is often at or above human levels, with systematic failures on cultural and numerical salience that culture prompts partially fix.

  5. The Incomplete Bridge: How AI Research (Mis)Engages with Psychology

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A citation-based study of 1,006 LLM papers finds psychology is increasingly cited, concentrated in psychometrics and neural mechanisms, and identifies repeated misapplications of Theory of Mind.

  6. Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task

    cs.CL 2025-07 conditional novelty 6.0 of 10

    In a new persuasion game, humans outperformed the LLM o1-preview when an opponent's preferences had to be inferred, while o1-preview outperformed humans when those preferences were disclosed.

  7. Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning with verifiable rewards makes a small LLM overfit theory-of-mind benchmarks, not acquire a generalizable theory of mind.

  8. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  9. From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Attention heads in multimodal LLMs linearly encode agents' beliefs, and steering those heads along probe-derived directions improves first- and second-order belief accuracy on the new GridToM benchmark.

  10. XToM: Exploring the Multilingual Theory of Mind for Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    XToM translates three English theory-of-mind benchmarks into Chinese, German, French, and Japanese with human quality control, and shows LLMs' belief reasoning is weaker and less consistent across languages than their...

  11. SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering

    cs.SE 2025-02 conditional novelty 6.0 of 10

    SyncBench quantifies agent out-of-sync recovery from 21 real repositories and finds state-of-the-art LLM agents succeed in under 34% of tasks and ask for help in under 5% of turns.

  12. VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

    cs.CL 2025-09 reject novelty 5.0 of 10

    A 65.6-hour movie-dialogue benchmark for spoken role-playing agents, with an evaluation framework whose main judge is also an evaluated model.

  13. LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

    cs.CL 2025-09 reject novelty 5.0 of 10

    LLMs can imitate mental-state annotation in team dialogue but systematically err on spatial reasoning and prosodic cues, per a six-dialogue CReST pilot.

  14. Referential ambiguity and clarification requests: comparing human and LLM behaviour

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Humans seldom ask clarification questions for referential ambiguity, while LLMs ask them more often, and reasoning prompts increase LLM question frequency and relevance.

  15. One Model, Two Minds: A Context-Gated Graph Learner that Recreates Human Biases

    cs.AI 2025-09 reject novelty 4.0 of 10

    A graph-based dual-process model with a learned gate claims to reproduce four cognitive biases in theory-of-mind tasks, but the bias effects are mostly learned from supervised labels rather than emergent.

  16. SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    SCOPE estimates a model's position bias with nonsense prompts, puts correct answers in disliked slots, and spreads similar distractors apart to cap lucky guessing.

  17. Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning

    cs.AI 2025-07 conditional novelty 4.0 of 10

    An LLM-augmented Bayesian inverse planning model, LAIP, generates hypotheses and action likelihoods, then uses Bayes' rule to infer agent preferences, outperforming LLM-only baselines.

  18. Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System

    cs.MA 2025-07 conditional novelty 4.0 of 10

    SynergyMAS combines a graph database with a Clingo logic solver, corrective RAG, and Theory of Mind prompts in a hierarchical multi-agent team, demonstrated on a Smart Home Energy Management case study.

  19. A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

    cs.LG 2025-06 unverdicted novelty 4.0 of 10

    This survey organizes foundation models, LLM agents, datasets, and tools in materials science into six task areas.

  20. When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

    cs.HC 2025-10 conditional novelty 3.0 of 10

    Researchers' claims of AI theory of mind are really about behavioral prediction, so AI evaluation should shift from isolated cognitive tests to human-AI interaction.

Pith tools