REVIEW 20 cited by
Evaluating Large Language Models in Theory of Mind Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The battery included 640 prompts spread across 40 diverse tasks, each one including a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. To solve a single task, a model needed to correctly answer 16 prompts across all eight scenarios. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. We explore the potential interpretation of these findings, including the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.
Forward citations
Cited by 20 Pith papers
-
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.
-
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
Credit assignment via LMM pairwise comparisons plus Bradley–Terry rank aggregation and potential-based shaping improves cooperative MARL under sparse rewards and dynamic agent counts.
-
AgentSociety 2: An Integrated Research Environment for Executable Social Science
An integrated LLM-agent environment runs social-science experiments from hypothesis to manuscript, reproducing several known human patterns while failing on others (implicit self-bias, free-riding decay, norm collapse).
-
Tacit Coordination of Large Language Models
Across 20+ open-source LLMs, tacit coordination in focal-point games is often at or above human levels, with systematic failures on cultural and numerical salience that culture prompts partially fix.
-
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
A citation-based study of 1,006 LLM papers finds psychology is increasingly cited, concentrated in psychometrics and neural mechanisms, and identifies repeated misapplications of Theory of Mind.
-
Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task
In a new persuasion game, humans outperformed the LLM o1-preview when an opponent's preferences had to be inferred, while o1-preview outperformed humans when those preferences were disclosed.
-
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
Reinforcement learning with verifiable rewards makes a small LLM overfit theory-of-mind benchmarks, not acquire a generalizable theory of mind.
-
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.
-
From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
Attention heads in multimodal LLMs linearly encode agents' beliefs, and steering those heads along probe-derived directions improves first- and second-order belief accuracy on the new GridToM benchmark.
-
XToM: Exploring the Multilingual Theory of Mind for Large Language Models
XToM translates three English theory-of-mind benchmarks into Chinese, German, French, and Japanese with human quality control, and shows LLMs' belief reasoning is weaker and less consistent across languages than their...
-
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
SyncBench quantifies agent out-of-sync recovery from 21 real repositories and finds state-of-the-art LLM agents succeed in under 34% of tasks and ask for help in under 5% of turns.
-
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
A 65.6-hour movie-dialogue benchmark for spoken role-playing agents, with an evaluation framework whose main judge is also an evaluated model.
-
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
LLMs can imitate mental-state annotation in team dialogue but systematically err on spatial reasoning and prosodic cues, per a six-dialogue CReST pilot.
-
Referential ambiguity and clarification requests: comparing human and LLM behaviour
Humans seldom ask clarification questions for referential ambiguity, while LLMs ask them more often, and reasoning prompts increase LLM question frequency and relevance.
-
One Model, Two Minds: A Context-Gated Graph Learner that Recreates Human Biases
A graph-based dual-process model with a learned gate claims to reproduce four cognitive biases in theory-of-mind tasks, but the bias effects are mostly learned from supervised labels rather than emergent.
-
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
SCOPE estimates a model's position bias with nonsense prompts, puts correct answers in disliked slots, and spreads similar distractors apart to cap lucky guessing.
-
Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning
An LLM-augmented Bayesian inverse planning model, LAIP, generates hypotheses and action likelihoods, then uses Bayes' rule to infer agent preferences, outperforming LLM-only baselines.
-
Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System
SynergyMAS combines a graph database with a Clingo logic solver, corrective RAG, and Theory of Mind prompts in a hierarchical multi-agent team, demonstrated on a Smart Home Energy Management case study.
-
A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
This survey organizes foundation models, LLM agents, datasets, and tools in materials science into six task areas.
-
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?
Researchers' claims of AI theory of mind are really about behavioral prediction, so AI evaluation should shift from isolated cognitive tests to human-AI interaction.
Discussion (0). Continue with ORCID to comment.