REVIEW 10 cited by
ToMBench: Benchmarking Theory of Mind in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Theory of Mind (ToM) is the cognitive capability to perceive and ascribe mental states to oneself and others. Recent research has sparked a debate over whether large language models (LLMs) exhibit a form of ToM. However, existing ToM evaluations are hindered by challenges such as constrained scope, subjective judgment, and unintended contamination, yielding inadequate assessments. To address this gap, we introduce ToMBench with three key characteristics: a systematic evaluation framework encompassing 8 tasks and 31 abilities in social cognition, a multiple-choice question format to support automated and unbiased evaluation, and a build-from-scratch bilingual inventory to strictly avoid data leakage. Based on ToMBench, we conduct extensive experiments to evaluate the ToM performance of 10 popular LLMs across tasks and abilities. We find that even the most advanced LLMs like GPT-4 lag behind human performance by over 10% points, indicating that LLMs have not achieved a human-level theory of mind yet. Our aim with ToMBench is to enable an efficient and effective evaluation of LLMs' ToM capabilities, thereby facilitating the development of LLMs with inherent social intelligence.
Forward citations
Cited by 10 Pith papers
-
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.
-
Mental World Modeling
Coupling physical and mental state in a world model, with target-specific observations and joint transitions, is necessary to predict human decisions across eight LLM backends on a process-annotated benchmark.
-
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
Unintentionally bad contexts (user suggestions or prior wrong assistant answers) cause LLMs to repeat errors, lose diversity, and flip stances, worsening with turns; RLVR on synthetic errors recovers 43–60% of the drop.
-
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.
-
XToM: Exploring the Multilingual Theory of Mind for Large Language Models
XToM translates three English theory-of-mind benchmarks into Chinese, German, French, and Japanese with human quality control, and shows LLMs' belief reasoning is weaker and less consistent across languages than their...
-
SocialEval: Evaluating Social Intelligence of Large Language Models
SocialEval is a 153-tree bilingual benchmark that evaluates LLM social intelligence through goal outcomes and interpersonal ability choices, finding LLMs below humans and biased toward prosocial behavior.
-
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue
LLMs can imitate mental-state annotation in team dialogue but systematically err on spatial reasoning and prosodic cues, per a six-dialogue CReST pilot.
-
H2HTalk: Evaluating Large Language Models as Emotional Companion
H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.
-
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
A new benchmark, UniToMBench, is proposed for evaluating Theory of Mind in LLMs, but its evaluation results are mixed and do not substantiate the claimed improvements.
-
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?
Researchers' claims of AI theory of mind are really about behavioral prediction, so AI evaluation should shift from isolated cognitive tests to human-AI interaction.
Discussion (0). Continue with ORCID to comment.