REVIEW 19 cited by
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Intuitive psychology is a pillar of common-sense reasoning. The replication of this reasoning in machine intelligence is an important stepping-stone on the way to human-like artificial intelligence. Several recent tasks and benchmarks for examining this reasoning in Large-Large Models have focused in particular on belief attribution in Theory-of-Mind tasks. These tasks have shown both successes and failures. We consider in particular a recent purported success case, and show that small variations that maintain the principles of ToM turn the results on their head. We argue that in general, the zero-hypothesis for model evaluation in intuitive psychology should be skeptical, and that outlying failure cases should outweigh average success rates. We also consider what possible future successes on Theory-of-Mind tasks by more powerful LLMs would mean for ToM tasks with people.
Forward citations
Cited by 19 Pith papers
-
Belief-reality separation lives in routing over a shared value slot in language models
Belief–reality separation in LMs lives in dissociated query-position routers over a frame-agnostic value slot filled by asserted binding or visibility-gated lookback.
-
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.
-
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.
-
Mental World Modeling
Coupling physical and mental state in a world model, with target-specific observations and joint transitions, is necessary to predict human decisions across eight LLM backends on a process-annotated benchmark.
-
Perceived AGI: Believability as Dimensional Completeness, Not Capability
A conversational agent's believability depends less on capability than on expressing four first-person stances — time, truth, entropy, love — that users read as evidence of a mind.
-
Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments
LLM agent groups show the human-style network-efficiency effect only when first-round choices are randomized; default agents start at the grid center and network topology has no measurable effect.
-
The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
Adding a structured list of six unknowable aspects of a user's life to an LLM prompt reduces sycophantic, harmful, and hallucinated advice in synthetic tests across five model families.
-
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
A DAG-based causal model specifies when theory of mind should engage in conflict — under information asymmetry, low accessible tractability, and perceived sophistication gaps — instead of treating mentalizing as always-on.
-
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism
ToM-U specifies a graph-based mechanism for epistemic state inference that derives belief states from behavior using LEWMs, bounded recursion, and a residue function for mentalizing failures.
-
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
Reinforcement learning with verifiable rewards makes a small LLM overfit theory-of-mind benchmarks, not acquire a generalizable theory of mind.
-
Strategy Adaptation in Large Language Model Werewolf Agents
Dynamically switching between Support and Attack strategies based on role estimates raises win rates for Werewolf LLM agents, with mixed effects for Villagers.
-
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.
-
From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
Attention heads in multimodal LLMs linearly encode agents' beliefs, and steering those heads along probe-derived directions improves first- and second-order belief accuracy on the new GridToM benchmark.
-
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
LLMs show limited, task-dependent accuracy at choosing correct epistemic modals and attitude verbs in controlled stories, with better performance on necessity and fact statements than on possibility and belief statements.
-
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
The DynToM benchmark shows ten LLMs average 33.0% accuracy versus 77.7% for humans, and models lose the most accuracy on questions about mental-state changes across scenarios.
-
Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?
LLMs fine-tuned on paraphrases of fictional facts can recall the paraphrases but cannot answer questions about who did what in those facts, suggesting memorization without robust scenario-level understanding.
-
Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning
An LLM-augmented Bayesian inverse planning model, LAIP, generates hypotheses and action likelihoods, then uses Bayes' rule to infer agent preferences, outperforming LLM-only baselines.
-
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
A new benchmark, UniToMBench, is proposed for evaluating Theory of Mind in LLMs, but its evaluation results are mixed and do not substantiate the claimed improvements.
-
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation
A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.
Discussion (0). Sign in to comment.