Pith. sign in

REVIEW 3 cited by

PHAnToM: Persona-based Prompting Has An Effect on Theory-of-Mind Reasoning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02246 v3 pith:JMU2C23U submitted 2024-03-04 cs.CL

classification cs.CL
keywords reasoningpromptingdifferencesllmsbeenlanguagemodelsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The use of LLMs in natural language reasoning has shown mixed results, sometimes rivaling or even surpassing human performance in simpler classification tasks while struggling with social-cognitive reasoning, a domain where humans naturally excel. These differences have been attributed to many factors, such as variations in prompting and the specific LLMs used. However, no reasons appear conclusive, and no clear mechanisms have been established in prior work. In this study, we empirically evaluate how role-playing prompting influences Theory-of-Mind (ToM) reasoning capabilities. Grounding our rsearch in psychological theory, we propose the mechanism that, beyond the inherent variance in the complexity of reasoning tasks, performance differences arise because of socially-motivated prompting differences. In an era where prompt engineering with role-play is a typical approach to adapt LLMs to new contexts, our research advocates caution as models that adopt specific personas might potentially result in errors in social-cognitive reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PeerPrism: Peer Evaluation Expertise vs Review-writing AI

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    PeerPrism benchmark demonstrates that state-of-the-art LLM detectors conflate surface text style with intellectual contribution and fail on hybrid human-AI peer reviews.

  2. Performance Analysis and Optimization for Laser-Phase-Noise based Quantum Random Number Generation

    quant-ph 2026-04 unverdicted novelty 6.0 of 10

    A comprehensive physical model predicts power spectrum and probability distribution in laser phase noise QRNG, enabling quantitative optimization of generation rate and quantum min-entropy.

  3. Performance Analysis and Optimization for Laser-Phase-Noise based Quantum Random Number Generation

    quant-ph 2026-04 unverdicted novelty 4.0 of 10

    A validated physical model predicts power spectrum and raw-data distributions for laser-phase-noise QRNGs, enabling quantitative rate optimization and proactive photonic-integrated design.

Pith tools