Pith. sign in

REVIEW 19 cited by

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.08399 v5 pith:ABTUNJ65 submitted 2023-02-16 cs.AI cs.CL

classification cs.AIcs.CL
keywords tasksreasoningtheory-of-mindconsiderintelligenceintuitivemodelsparticular
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intuitive psychology is a pillar of common-sense reasoning. The replication of this reasoning in machine intelligence is an important stepping-stone on the way to human-like artificial intelligence. Several recent tasks and benchmarks for examining this reasoning in Large-Large Models have focused in particular on belief attribution in Theory-of-Mind tasks. These tasks have shown both successes and failures. We consider in particular a recent purported success case, and show that small variations that maintain the principles of ToM turn the results on their head. We argue that in general, the zero-hypothesis for model evaluation in intuitive psychology should be skeptical, and that outlying failure cases should outweigh average success rates. We also consider what possible future successes on Theory-of-Mind tasks by more powerful LLMs would mean for ToM tasks with people.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 79 citations worldwide. Full citation record

  1. Belief-reality separation lives in routing over a shared value slot in language models

    cs.CL 2026-07 conditional novelty 7.5 of 10

    Belief–reality separation in LMs lives in dissociated query-position routers over a frame-agnostic value slot filled by asserted binding or visibility-gated lookback.

  2. MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.

  3. Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.

  4. Mental World Modeling

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Coupling physical and mental state in a world model, with target-specific observations and joint transitions, is necessary to predict human decisions across eight LLM backends on a process-annotated benchmark.

  5. Perceived AGI: Believability as Dimensional Completeness, Not Capability

    cs.HC 2026-07 conditional novelty 6.0 of 10

    A conversational agent's believability depends less on capability than on expressing four first-person stances — time, truth, entropy, love — that users read as evidence of a mind.

  6. Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments

    cs.AI 2026-07 accept novelty 6.0 of 10

    LLM agent groups show the human-style network-efficiency effect only when first-round choices are randomized; default agents start at the grid center and network topology has no measurable effect.

  7. The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Adding a structured list of six unknowable aspects of a user's life to an LLM prompt reduces sycophantic, harmful, and hallucinated advice in synthetic tests across five model families.

  8. A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

    cs.AI 2026-06 conditional novelty 6.0 of 10

    A DAG-based causal model specifies when theory of mind should engage in conflict — under information asymmetry, low accessible tractability, and perceived sophistication gaps — instead of treating mentalizing as always-on.

  9. The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    ToM-U specifies a graph-based mechanism for epistemic state inference that derives belief states from behavior using LEWMs, bounded recursion, and a residue function for mentalizing failures.

  10. Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning with verifiable rewards makes a small LLM overfit theory-of-mind benchmarks, not acquire a generalizable theory of mind.

  11. Strategy Adaptation in Large Language Model Werewolf Agents

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Dynamically switching between Support and Attack strategies based on role estimates raises win rates for Werewolf LLM agents, with mixed effects for Villagers.

  12. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  13. From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Attention heads in multimodal LLMs linearly encode agents' beliefs, and steering those heads along probe-derived directions improves first- and second-order belief accuracy on the new GridToM benchmark.

  14. Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs show limited, task-dependent accuracy at choosing correct epistemic modals and attitude verbs in controlled stories, with better performance on necessity and fact statements than on possibility and belief statements.

  15. Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The DynToM benchmark shows ten LLMs average 33.0% accuracy versus 77.7% for humans, and models lose the most accuracy on questions about mental-state changes across scenarios.

  16. Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?

    cs.CL 2025-09 conditional novelty 5.0 of 10

    LLMs fine-tuned on paraphrases of fictional facts can recall the paraphrases but cannot answer questions about who did what in those facts, suggesting memorization without robust scenario-level understanding.

  17. Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning

    cs.AI 2025-07 conditional novelty 4.0 of 10

    An LLM-augmented Bayesian inverse planning model, LAIP, generates hypotheses and action likelihoods, then uses Bayes' rule to infer agent preferences, outperforming LLM-only baselines.

  18. UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs

    cs.CL 2025-06 reject novelty 4.0 of 10

    A new benchmark, UniToMBench, is proposed for evaluating Theory of Mind in LLMs, but its evaluation results are mixed and do not substantiate the claimed improvements.

  19. Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.

Pith tools