Pith. sign in

REVIEW 4 cited by

Boosting Theory-of-Mind Performance in Large Language Models via Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11490 v3 pith:P7UU72CP submitted 2023-04-22 cs.AI cs.CL

classification cs.AIcs.CL
keywords accuracylearningreasoninggpt-4in-contextllmsmodelsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) excel in many tasks in 2023, but they still face challenges in complex reasoning. Theory-of-mind (ToM) tasks, which require understanding agents' beliefs, goals, and mental states, are essential for common-sense reasoning involving humans, making it crucial to enhance LLM performance in this area. This study measures the ToM performance of GPT-4 and three GPT-3.5 variants (Davinci-2, Davinci-3, GPT-3.5-Turbo), and investigates the effectiveness of in-context learning in improving their ToM comprehension. We evaluated prompts featuring two-shot chain of thought reasoning and step-by-step thinking instructions. We found that LLMs trained with Reinforcement Learning from Human Feedback (RLHF) (all models excluding Davinci-2) improved their ToM accuracy via in-context learning. GPT-4 performed best in zero-shot settings, reaching nearly 80% ToM accuracy, but still fell short of the 87% human accuracy on the test set. However, when supplied with prompts for in-context learning, all RLHF-trained LLMs exceeded 80% ToM accuracy, with GPT-4 reaching 100%. These results demonstrate that appropriate prompting enhances LLM ToM reasoning, and they underscore the context-dependent nature of LLM cognitive capacities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Decompose-ToM recursively decomposes theory-of-mind questions into perspective simulations and knowledge-access checks, improving accuracy on Hi-ToM but only matching SimToM on FANToM.

  2. Detecting Conversational Mental Manipulation with Intent-Aware Prompting

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Adding per-speaker intent summaries to an LLM prompt reduces false negatives in mental manipulation detection by 30.5% versus zero-shot prompting on the MentalManip dataset.

  3. TARS: A Theory-of-Mind Agent for Personalized In-IDE Code Comprehension

    cs.SE 2026-07 conditional novelty 5.0 of 10

    An in-IDE Theory-of-Mind agent produced suggestive, non-significant speed gains and self-reported personalization benefits in an 18-developer study.

  4. UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs

    cs.CL 2025-06 reject novelty 4.0 of 10

    A new benchmark, UniToMBench, is proposed for evaluating Theory of Mind in LLMs, but its evaluation results are mixed and do not substantiate the claimed improvements.

Pith tools