REVIEW 4 cited by
Boosting Theory-of-Mind Performance in Large Language Models via Prompting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) excel in many tasks in 2023, but they still face challenges in complex reasoning. Theory-of-mind (ToM) tasks, which require understanding agents' beliefs, goals, and mental states, are essential for common-sense reasoning involving humans, making it crucial to enhance LLM performance in this area. This study measures the ToM performance of GPT-4 and three GPT-3.5 variants (Davinci-2, Davinci-3, GPT-3.5-Turbo), and investigates the effectiveness of in-context learning in improving their ToM comprehension. We evaluated prompts featuring two-shot chain of thought reasoning and step-by-step thinking instructions. We found that LLMs trained with Reinforcement Learning from Human Feedback (RLHF) (all models excluding Davinci-2) improved their ToM accuracy via in-context learning. GPT-4 performed best in zero-shot settings, reaching nearly 80% ToM accuracy, but still fell short of the 87% human accuracy on the test set. However, when supplied with prompts for in-context learning, all RLHF-trained LLMs exceeded 80% ToM accuracy, with GPT-4 reaching 100%. These results demonstrate that appropriate prompting enhances LLM ToM reasoning, and they underscore the context-dependent nature of LLM cognitive capacities.
Forward citations
Cited by 4 Pith papers
-
Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
Decompose-ToM recursively decomposes theory-of-mind questions into perspective simulations and knowledge-access checks, improving accuracy on Hi-ToM but only matching SimToM on FANToM.
-
Detecting Conversational Mental Manipulation with Intent-Aware Prompting
Adding per-speaker intent summaries to an LLM prompt reduces false negatives in mental manipulation detection by 30.5% versus zero-shot prompting on the MentalManip dataset.
-
TARS: A Theory-of-Mind Agent for Personalized In-IDE Code Comprehension
An in-IDE Theory-of-Mind agent produced suggestive, non-significant speed gains and self-reported personalization benefits in an 18-developer study.
-
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
A new benchmark, UniToMBench, is proposed for evaluating Theory of Mind in LLMs, but its evaluation results are mixed and do not substantiate the claimed improvements.
Discussion (0). Continue with ORCID to comment.