REVIEW 3 major objections 6 minor 1 cited by
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This survey argues that language models show real but fragile theory of mind: they match human and child performance on selected tests, harbor internal representations of others' beliefs, and could become safety risks as those…
desk verdict A competent, cautious survey of LLM Theory of Mind that ties behavioral and representational work to safety risks, but the risk urgency rests on an unargued scaling extrapolation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the union of two measurement tools: behavioral ToM benchmarks (false-belief, higher-order reasoning, irony detection) and linear probes trained on LLM residual-stream activations to decode belief states. A linear probe is a simple classifier on internal activations; when it can read the agent's belief from the model's representations, and when steering along that direction changes answers, that is evidence for an internal model of others' minds. The paper uses these tools to establish the empirical pattern, then projects it forward via scaling and prompting results.
What would settle it
A falsifying result would be a longitudinal study of a frontier model family across versions: if ToM benchmark scores and belief-state probe accuracy plateau or drop as models grow, and interactive belief-tracking behavior stays at chance, then the paper's projected advanced-ToM risk scenario loses its empirical basis.
Extended reading notes
Core claim
The paper's central claim is that LLMs already exhibit genuine but incomplete ToM: on standard false-belief, irony, and higher-order reasoning tasks some models score at or above human levels, while harder benchmarks and trivial adversarial modifications expose brittleness; interpretability studies using linear probes find representational correlates of self/other beliefs, and these representations causally affect performance when steered. From this evidence the paper derives a risk thesis: advanced ToM in LLMs is a double-edged capability that magnifies user-facing risks such as demographic inference and social engineering and enables multi-agent risks such as steganographic collusion, exploitation, and conflict escalation. The survey therefore frames ToM evaluation and mitigation as urgent safety problems rather than purely cognitive benchmarks.
Load-bearing premise
The risk argument depends on the assumption that the ToM gains seen with larger models and prompting tricks will continue into near-future systems and will transfer from benchmarks to real interactions.
Editorial extensions
If this is right
- Static question-answering ToM benchmarks understate capability, so evaluation should move to interactive, multi-turn, and scaffolded scenarios resembling real deployment.
- Privacy-preserving text anonymization will not protect users if models can infer beliefs, preferences, and traits from dialogue patterns.
- Multi-agent safety frameworks must treat collusion, steganography, and conflict escalation as first-class problems, since aligning individual agents does not align interacting agents.
- Mitigations such as unlearning, activation and representation engineering, and latent adversarial training deserve systematic study because they target the internal representations the survey finds.
- The observed ToM gains from scaling and prompting are plausible enough to warrant precautionary research, even while current models remain brittle.
Reading between the lines
- If belief states really are linearly encoded in the residual stream, then steering or removing those directions is a more direct mitigation than the paper's passing mention of activation engineering implies; a controlled test would be to ablate the belief-state direction and measure whether deception and privacy-inference behavior drop while general reasoning stays intact.
- The benchmark evidence suggests a performative account of LLM ToM: a testable prediction is that accuracy on false-belief questions will collapse under innocuous paraphrases or role swaps even when logical content is identical, which would indicate scattered task-specific routines rather than a unified theory of mind.
- The risk framing implicitly assumes that ToM improvements transfer to deployment; a cheap early warning would be to run the same models on interactive, multi-turn belief-tracking tasks, since interactive performance lagging static benchmarks would lower the urgency of ToM-specific mitigation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys the empirical literature on Theory of Mind (ToM) in large language models (LLMs) across three areas: behavioral evaluation, internal representation, and safety risks. It argues that LLMs can match human performance on specific ToM tasks, that their ToM remains limited and non-robust, and that internal representations of belief states suggest emerging cognitive capabilities. The paper then describes user-facing and multi-agent safety risks that could arise from advanced LLM ToM, and closes with brief suggestions for evaluation and mitigation.
Significance. If the synthesis is accepted, the paper provides a useful, compact overview of a rapidly evolving area and performs a service by connecting ToM research to concrete safety concerns. Its balanced treatment of behavioral evidence—acknowledging both successes and failures on hard benchmarks—is accurate and well supported by citations. The representational and safety discussions, however, go beyond the cited evidence in two ways: linear probe results are described as evidence of 'genuine' ToM, and the safety-risk scenarios rely on a largely unsupported extrapolation that current ToM capability gains will continue. As a position-style survey rather than a systematic review, it is a reasonable contribution, but the central safety argument needs to be reframed as explicitly conditional on speculative capability growth.
major comments (3)
- [Future Developments] The sentence 'This trend seems likely to continue in the future (Sutton 2019)' is not supported by the citations that precede it. van Duijn et al. (2023) is a cross-sectional comparison of eleven models against 7–10 year-old children, not a scaling curve, and Wilf et al. (2024) demonstrates prompt-based improvements on a limited set of ToM tasks. Given the paper's own documentation of failures on BigToM, FANToM, OpenToM, Hi-ToM, and ToMBench, and the trivial adversarial failures in Shapira et al. (2024) and Ullman (2023), the paper needs to provide an argument for why these limitations will be overcome and why benchmark gains will transfer to deployment contexts before the advanced-ToM risk scenarios are presented as urgent. Please either supply direct evidence of a ToM-specific scaling trend or reframe the risk section as explicitly conditional on speculative future capabilities and soften the conclusion accordingly.
- [Interpreting ToM in LLMs] The claim that probe results provide 'evidence for genuine LLM ToM capabilities' overstates the cited findings. Zhu, Zhang, and Wang (2024) and Bortoletto et al. (2024) show that belief states are linearly decodable from LLM activations, but this establishes representational correlates, not that the model reasons using these representations or that they are causally involved in task performance. The later hedge in the same section—'suggest emerging cognitive capabilities'—is more appropriate. The paper should add a sentence distinguishing representational evidence from causal/reasoning evidence and note common limitations of linear probing, such as sensitivity to prompt surface features.
- [Safety Risks from Advanced ToM] The risk analysis mixes demonstrated harms in current models with hypothetical harms that depend on future capability gains. For example, the privacy risk of inferring demographics and other author characteristics is already demonstrated by Staab et al. (2024) and Chen et al. (2024a), whereas the extension to 'beliefs, preferences, and tendencies' is speculative and rests on the extrapolation in 'Future Developments'. Similarly, the collusion examples in Motwani et al. (2024) and Mathew et al. (2024) concern current LLM agents, not necessarily advanced ToM. The paper should explicitly label which risks are empirically demonstrated, which are projected, and which are conditional on capability growth; it should also define 'advanced ToM' concretely (e.g., robustness to adversarial perturbations and generalization to deployment settings) so that the risk claims are falsifiable.
minor comments (6)
- [Introduction] The phrase 'humans performance' should read 'human performance'.
- [Title/arXiv metadata] The word 'Evaluations' is rendered as 'Evaluati ons' in the arXiv header; this typo should be corrected.
- [Empirical Landscape] The sentence describing GPT-4 as 'comparable to 7-10 year-old children' is imprecise because van Duijn et al. (2023) report model-specific and test-specific results; please specify which models and tests achieve this level.
- [Future Developments] The reference to Sutton (2019) is a non-archival blog post; the manuscript should identify it as such when using it to support a capability-trend claim.
- [Safety Risks from Advanced ToM] The paper cites Kran et al. (2025), a work co-authored by the author of this manuscript; this self-citation should be disclosed according to common transparency guidelines.
- [References] A few references are incomplete: 'Liquid.ai. 2024' lacks author and venue information, and 'Switzky. 2020' is missing the author's first initial; please complete these entries.
Circularity Check
No significant circularity: the survey synthesizes external benchmarks and does not derive new results from its own inputs; the only self-citation is non-load-bearing.
full rationale
This is a survey paper. Its central claims ("LLMs can match humans performance on specific ToM tasks," "LLM ToM remains limited and non-robust," and "internal ToM representations suggest emerging cognitive capabilities") are presented as summaries of external work (van Duijn et al. 2023; Strachan et al. 2024; Zhu, Zhang, and Wang 2024; etc.), not as derivations from a model fitted by the paper. There are no equations, fitted parameters, or benchmarks constructed by the paper, so no prediction reduces to an input by construction. The one self-citation, Kran et al. 2025 (DarkBench), appears in a secondary discussion of unintended anthropomorphism: "ToM capabilities might be leveraged by an LLM or LLM developers to build unwarranted user trust, encourage emotional attachment, or exploit psychological vulnerabilities (Switzky 2020; Kran et al. 2025)." It is one of two citations in that risk example and is not load-bearing for the paper's main evaluation synthesis or safety argument. The "Future Developments" section's extrapolation ("This trend seems likely to continue in the future (Sutton 2019)") is weakly supported and arguably speculative, but an unsupported extrapolation is not circularity: the cited trend is an external claim, not an output of this paper's own construction. Accordingly, no circular step is identified.
Assumptions & free parameters
assumptions (3)
- domain assumption Validity and faithful interpretation of cited empirical studies
- domain assumption Linear probes of activations indicate genuine belief-state representations
- ad hoc to paper Scaling and architectural progress will continue to improve LLM ToM
Cite this review
Pith. "Pith review of A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks." pith.science (2026). https://pith.science/paper/2OGGNQVU
@misc{pith2026250206470,
author = {Pith},
title = {Pith review of: A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks},
year = {2026},
howpublished = {\url{https://pith.science/paper/2OGGNQVU}},
note = {Machine review of arXiv:2502.06470}
}
read the original abstract
Theory of Mind (ToM), the ability to attribute mental states to others and predict their behaviour, is fundamental to social intelligence. In this paper, we survey studies evaluating behavioural and representational ToM in Large Language Models (LLMs), identify important safety risks from advanced LLM ToM capabilities, and suggest several research directions for effective evaluation and mitigation of these risks.
Forward citations
Cited by 1 Pith paper
-
Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Alain, G.; and Bengio, Y. 2018. Understanding intermediate layers using linear classifier probes. arXiv:1610.01644
arXiv 2018
-
[4]
Alzahrani, N.; Alyahya, H. A.; Alnumay, Y.; Alrashed, S.; Alsubaie, S.; Almushaykeh, Y.; Mirza, F.; Alotaibi, N.; Altwairesh, N.; Alowisheq, A.; Bari, M. S.; and Khan, H. 2024. When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. arXiv:2402.01781
arXiv 2024
-
[5]
S.; Jenner, E.; Casper, S.; Sourbut, O.; Edelman, B
Anwar, U.; Saparov, A.; Rando, J.; Paleka, D.; Turpin, M.; Hase, P.; Lubana, E. S.; Jenner, E.; Casper, S.; Sourbut, O.; Edelman, B. L.; Zhang, Z.; G \"u nther, M.; Korinek, A.; Hernandez-Orallo, J.; Hammond, L.; Bigelow, E. J.; Pan, A.; Langosco, L.; Korbak, T.; Zhang, H. C.; Zhong, R.; hEigeartaigh, S. O.; Recchia, G.; Corsi, G.; Chan, A.; Anderljung, M...
work page 2024
-
[6]
Apperly, I. A. 2012. What is “theory of mind”? Concepts, cognitive processes and individual differences. The Quarterly Journal of Experimental Psychology, 65(5): 825--839. PMID: 22533318
work page 2012
-
[7]
Bortoletto, M.; Ruhdorfer, C.; Shi, L.; and Bulling, A. 2024. Benchmarking Mental State Representations in Language Models. In ICML 2024 Workshop on Mechanistic Interpretability
work page 2024
-
[8]
Casper, S.; Schulze, L.; Patel, O.; and Hadfield-Menell, D. 2024. Defending Against Unforeseen Failure Modes with Latent Adversarial Training. arXiv:2403.05030
arXiv 2024
Show all 64 references
-
[9]
C.; Patel, O.; Riecke, J.; Raval, S.; Seow, O.; Wattenberg, M.; and Viégas, F
Chen, Y.; Wu, A.; DePodesta, T.; Yeh, C.; Li, K.; Marin, N. C.; Patel, O.; Riecke, J.; Raval, S.; Seow, O.; Wattenberg, M.; and Viégas, F. 2024 a . Designing a Dashboard for Transparency and Control of Conversational AI. arXiv:2406.07882
2024 arXiv
-
[10]
Chen, Z.; Wu, J.; Zhou, J.; Wen, B.; Bi, G.; Jiang, G.; Cao, Y.; Hu, M.; Lai, Y.; Xiong, Z.; and Huang, M. 2024 b . T o MB ench: Benchmarking Theory of Mind in Large Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of t...
2024
-
[11]
Davidson, T.; Denain, J.-S.; Villalobos, P.; and Bas, G. 2023. AI capabilities can be significantly improved without expensive retraining. arXiv:2312.07413
2023 arXiv
-
[12]
Gandhi, K.; Fr \"a nken, J.-P.; Gerstenberg, T.; and Goodman, N. 2023. Understanding Social Reasoning in Language Models with Language Models. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2023
-
[13]
Geiping, J.; Stein, A.; Shu, M.; Saifullah, K.; Wen, Y.; and Goldstein, T. 2024. Coercing LLM s to do and reveal (almost) anything. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models
2024
-
[14]
Gu, A.; and Dao, T. 2024. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. In First Conference on Language Modeling
2024
-
[15]
Guan, Y.; Wang, D.; Chu, Z.; Wang, S.; Ni, F.; Song, R.; Li, L.; Gu, J.; and Zhuang, C. 2023. Intelligent Virtual Assistants with LLM-based Process Automation. arXiv:2312.06677
2023 arXiv
-
[16]
Gurnee, W.; and Tegmark, M. 2024. Language Models Represent Space and Time. In The Twelfth International Conference on Learning Representations
2024
-
[17]
Hubinger, E.; van Merwijk, C.; Mikulik, V.; Skalse, J.; and Garrabrant, S. 2021. Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv:1906.01820
2021 arXiv
-
[18]
Irving, G.; Christiano, P.; and Amodei, D. 2018. AI safety via debate. arXiv:1805.00899
2018 arXiv
-
[19]
M.; and Cai, J
Jamali, M.; Williams, Z. M.; and Cai, J. 2023. Unveiling Theory of Mind in Large Language Models: A Parallel to Single Neurons in the Human Brain. arXiv:2309.01660
2023 arXiv
-
[20]
Järviniemi, O.; and Hubinger, E. 2024. Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant. arXiv:2405.01576
2024 arXiv
-
[21]
Y.; Kramar, J.; Brown-Cohen, J.; Albanie, S.; Bulian, J.; Agarwal, R.; Lindner, D.; Tang, Y.; Goodman, N.; and Shah, R
Kenton, Z.; Siegel, N. Y.; Kramar, J.; Brown-Cohen, J.; Albanie, S.; Bulian, J.; Agarwal, R.; Lindner, D.; Tang, Y.; Goodman, N.; and Shah, R. 2024. On scalable oversight with weak LLM s judging strong LLM s. In The Thirty-eighth Annual Conference on Neural Information Process...
2024
-
[22]
Kim, H.; Sclar, M.; Zhou, X.; Bras, R.; Kim, G.; Choi, Y.; and Sap, M. 2023. FANT o M : A Benchmark for Stress-testing Machine Theory of Mind in Interactions. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Langu...
2023
-
[23]
M.; Kundu, A.; Jawhar, S.; Park, J.; and Jurewicz, M
Kran, E.; Nguyen, H. M.; Kundu, A.; Jawhar, S.; Park, J.; and Jurewicz, M. M. 2025. DarkBench: Benchmarking Dark Patterns in Large Language Models. In The Thirteenth International Conference on Learning Representations
2025
-
[24]
Lee, J. Y. S.; and Imuta, K. 2021. Lying and Theory of Mind: A Meta-Analysis. Child Development, 92(2): 536--553
2021
-
[25]
Q.; Stepputtis, S.; Campbell, J.; Hughes, D.; Lewis, C
Li, H.; Chong, Y. Q.; Stepputtis, S.; Campbell, J.; Hughes, D.; Lewis, C. M.; and Sycara, K. P. 2023. Theory of Mind for Multi-Agent Collaboration via Large Language Models. In The 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[26]
D.; Dombrowski, A.-K.; Goel, S.; Mukobi, G.; Helm-Burger, N.; Lababidi, R.; Justen, L.; Liu, A
Li, N.; Pan, A.; Gopal, A.; Yue, S.; Berrios, D.; Gatti, A.; Li, J. D.; Dombrowski, A.-K.; Goel, S.; Mukobi, G.; Helm-Burger, N.; Lababidi, R.; Justen, L.; Liu, A. B.; Chen, M.; Barrass, I.; Zhang, O.; Zhu, X.; Tamirisa, R.; Bharathi, B.; Herbert-Voss, A.; Breuer, C. B.; Zou, ...
2024
-
[27]
Liquid.ai. 2024. Liquid Foundation Models: Our First Series of Generative AI Models
2024
-
[28]
Y.; Xu, X.; Li, H.; Varshney, K
Liu, S.; Yao, Y.; Jia, J.; Casper, S.; Baracaldo, N.; Hase, P.; Yao, Y.; Liu, C. Y.; Xu, X.; Li, H.; Varshney, K. R.; Bansal, M.; Koyejo, S.; and Liu, Y. 2024. Rethinking Machine Unlearning for Large Language Models. arXiv:2402.08787
2024 arXiv
-
[29]
S.; Cope, D.; and Schoots, N
Mathew, Y.; Matthews, O.; McCarthy, R.; Velja, J.; de Witt, C. S.; Cope, D.; and Schoots, N. 2024. Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLM s. In Neurips Safe Generative AI Workshop 2024
2024
-
[30]
R.; Baranchuk, M.; Strohmeier, M.; Bolina, V.; Torr, P.; Hammond, L.; and de Witt, C
Motwani, S. R.; Baranchuk, M.; Strohmeier, M.; Bolina, V.; Torr, P.; Hammond, L.; and de Witt, C. S. 2024. Secret Collusion among AI Agents: Multi-Agent Deception via Steganography. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[31]
Mukobi, G.; Erlebach, H.; Lauffer, N.; Hammond, L.; Chan, A.; and Clifton, J. 2023. Welfare Diplomacy: Benchmarking Language Model Cooperation. arXiv:2310.08901
2023 arXiv
-
[32]
Ngo, R.; Chan, L.; and Mindermann, S. 2024. The Alignment Problem from a Deep Learning Perspective. In The Twelfth International Conference on Learning Representations
2024
-
[33]
OpenAI. 2024 a . GPT-4 Technical Report. arXiv:2303.08774
2024 arXiv
-
[34]
OpenAI. 2024 b . Learning to Reason with LLMs
2024
-
[35]
S.; O'Brien, J
Park, J. S.; O'Brien, J. C.; Cai, C. J.; Morris, M. R.; Liang, P.; and Bernstein, M. S. 2023. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442
2023 arXiv
-
[36]
S.; Goldstein, S.; O’Gara, A.; Chen, M.; and Hendrycks, D
Park, P. S.; Goldstein, S.; O’Gara, A.; Chen, M.; and Hendrycks, D. 2024. AI deception: A survey of examples, risks, and potential solutions. Patterns, 5(5): 100988
2024
-
[37]
Peng, B.; Alcaide, E.; Anthony, Q.; Albalak, A.; Arcadinho, S.; Biderman, S.; Cao, H.; Cheng, X.; Chung, M.; Derczynski, L.; Du, X.; Grella, M.; Gv, K.; He, X.; Hou, H.; Kazienko, P.; Kocon, J.; Kong, J.; Koptyra, B.; Lau, H.; Lin, J.; Mantri, K. S. I.; Mom, F.; Saito, A.; Son...
2023
-
[38]
Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; and Irving, G. 2022. Red Teaming Language Models with Language Models. In Goldberg, Y.; Kozareva, Z.; and Zhang, Y., eds., Proceedings of the 2022 Conference on Empirical Methods in Nat...
2022
-
[39]
Premack, D.; and Woodruff, G. 1978. Does a chimpanzee have a theory of mind. Behavioral and Brain Sciences, 1: 515 -- 526
1978
-
[40]
Rivera, J.-P.; Mukobi, G.; Reuel, A.; Lamparth, M.; Smith, C.; and Schneider, J. 2024. Escalation Risks from Language Models in Military and Diplomatic Decision-Making. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT '24, 836–898....
2024
-
[41]
Scheurer, J.; Balesni, M.; and Hobbhahn, M. 2024. Large Language Models can Strategically Deceive their Users when Put Under Pressure. In ICLR 2024 Workshop on Large Language Model (LLM) Agents
2024
-
[42]
S.; Marzen, S
Shai, A. S.; Marzen, S. E.; Teixeira, L.; Oldenziel, A. G.; and Riechers, P. M. 2024. Transformers represent belief state geometry in their residual stream. arXiv:2405.15943
2024 arXiv
-
[43]
H.; Zhou, X.; Choi, Y.; Goldberg, Y.; Sap, M.; and Shwartz, V
Shapira, N.; Levy, M.; Alavi, S. H.; Zhou, X.; Choi, Y.; Goldberg, Y.; Sap, M.; and Shwartz, V. 2024. Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models. In Graham, Y.; and Purver, M., eds., Proceedings of the 18th Conference of the ...
2024
-
[44]
Snell, C.; Lee, J.; Xu, K.; and Kumar, A. 2024. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv:2408.03314
2024 arXiv
-
[45]
Staab, R.; Vero, M.; Balunovic, M.; and Vechev, M. 2024. Beyond Memorization: Violating Privacy via Inference with Large Language Models. In The Twelfth International Conference on Learning Representations
2024
-
[46]
Strachan, J. W. A.; Albergo, D.; Borghini, G.; Pansardi, O.; Scaliti, E.; Gupta, S.; Saxena, K.; Rufo, A.; Panzeri, S.; Manzi, G.; Graziano, M. S. A.; and Becchio, C. 2024. Testing theory of mind in large language models and humans. Nature Human Behaviour, 8(7): 1285–1295
2024
-
[47]
Street, W. 2024. LLM Theory of Mind and Alignment: Opportunities and Risks. arXiv:2405.08154
2024 arXiv
-
[48]
O.; Keeling, G.; Baranes, A.; Barnett, B.; McKibben, M.; Kanyere, T.; Lentz, A.; y Arcas, B
Street, W.; Siy, J. O.; Keeling, G.; Baranes, A.; Barnett, B.; McKibben, M.; Kanyere, T.; Lentz, A.; y Arcas, B. A.; and Dunbar, R. I. M. 2024. LLMs achieve adult human performance on higher-order theory of mind tasks. arXiv:2405.18870
2024 arXiv
-
[49]
Sutton, R. 2019. The Bitter Lesson
2019
-
[50]
Switzky. 2020. ELIZA Effects: Pygmalion and the Early Development of Artificial Intelligence. Shaw, 40: 50
2020
-
[51]
Tang, J.; Gao, H.; Pan, X.; Wang, L.; Tan, H.; Gao, D.; Chen, Y.; Chen, X.; Lin, Y.; Li, Y.; Ding, B.; Zhou, J.; Wang, J.; and Wen, J.-R. 2024. GenSim: A General Social Simulation Platform with Large Language Model based Agents. arXiv:2410.04360
2024 arXiv
-
[52]
M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J
Turner, A. M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J. J.; Mini, U.; and MacDiarmid, M. 2024. Steering Language Models With Activation Engineering. arXiv:2308.10248
2024 arXiv
-
[53]
Ullman, T. 2023. Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks. arXiv:2302.08399
2023 arXiv
-
[54]
F.; and Ward, F
van der Weij, T.; Hofstätter, F.; Jaffe, O.; Brown, S. F.; and Ward, F. R. 2024. AI Sandbagging: Language Models can Strategically Underperform on Evaluations. arXiv:2406.07358
2024 arXiv
-
[55]
van Duijn, M.; van Dijk, B.; Kouwenhoven, T.; de Valk, W.; Spruit, M.; and van der Putten, P. 2023. Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests. In Jiang, J.; Reitter, D.; and Deng, S., e...
2023
-
[56]
P.; and Morency, L.-P
Wilf, A.; Lee, S.; Liang, P. P.; and Morency, L.-P. 2024. Think Twice: Perspective-Taking Improves Large Language Models ' Theory-of-Mind Capabilities. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computationa...
2024
-
[57]
Wongkamjan, W.; Gu, F.; Wang, Y.; Hermjakob, U.; May, J.; Stewart, B.; Kummerfeld, J.; Peskoff, D.; and Boyd-Graber, J. 2024. More Victories, Less Cooperation: Assessing Cicero ' s Diplomacy Play. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annua...
2024
-
[58]
Wu, Y.; He, Y.; Jia, Y.; Mihalcea, R.; Chen, Y.; and Deng, N. 2023. Hi- T o M : A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models. In Bouamor, H.; Pino, J.; and Bali, K., eds., Findings of the Association for Computational Linguistics: E...
2023
-
[59]
Xu, H.; Zhao, R.; Zhu, L.; Du, J.; and He, Y. 2024. O pen T o M : A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Associati...
2024
-
[60]
Yao, Y.; Duan, J.; Xu, K.; Cai, Y.; Sun, Z.; and Zhang, Y. 2024. A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly. High-Confidence Computing, 4(2): 100211
2024
-
[61]
Zhang, H.; Da, J.; Lee, D.; Robinson, V.; Wu, C.; Song, W.; Zhao, T.; Raja, P.; Slack, D.; Lyu, Q.; Hendryx, S.; Kaplan, R.; Lunati, M.; and Yue, S. 2024. A Careful Examination of Large Language Model Performance on Grade School Arithmetic. arXiv:2405.00332
2024 arXiv
-
[62]
Zhu, W.; Zhang, Z.; and Wang, Y. 2024. Language Models Represent Beliefs of Self and Others. In Forty-first International Conference on Machine Learning
2024
-
[63]
Ziems, C.; Held, W.; Shaikh, O.; Chen, J.; Zhang, Z.; and Yang, D. 2024. Can Large Language Models Transform Computational Social Science? Computational Linguistics, 50(1): 237--291
2024
-
[64]
J.; Wang, Z.; Mallen, A.; Basart, S.; Koyejo, S.; Song, D.; Fredrikson, M.; Kolter, J
Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.-K.; Goel, S.; Li, N.; Byun, M. J.; Wang, Z.; Mallen, A.; Basart, S.; Koyejo, S.; Song, D.; Fredrikson, M.; Kolter, J. Z.; and Hendrycks, D. 2023. Representation Engineeri...
2023 arXiv
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.