Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2302.02083.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:11:57.606122Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
176
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 7ebc297f-623f-474e-83f2-d993cace0173 · inbound
Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Evaluating Large Language Models in Theory of Mind Tasks
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83bf78cd-26e0-44dc-b9d9-1f57f27beeb2 · inbound
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Evaluating Large Language Models in Theory of Mind Tasks
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 60fe1ccc-95a3-41cf-b7f5-c3cfd7f4ada9 · inbound
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering Evaluating Large Language Models in Theory of Mind Tasks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439ce8ad-d210-49df-999e-122040f2c452 · inbound
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments Evaluating Large Language Models in Theory of Mind Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b8e7fb44-2314-44bf-b0f9-3eba03df5ff4 · inbound
XToM: Exploring the Multilingual Theory of Mind for Large Language Models Evaluating Large Language Models in Theory of Mind Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818afb4d-91ae-48c4-a0b6-0061a4073722 · inbound
From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Evaluating Large Language Models in Theory of Mind Tasks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5176ce3-a507-4e6e-93cc-7aa339c8e6fd · inbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db051cc-f060-4333-ac85-9e32efa28e73 · inbound
A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools Evaluating Large Language Models in Theory of Mind Tasks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a69d7a8-8997-4c9d-955a-dc06a9ba2835 · inbound
Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System Evaluating Large Language Models in Theory of Mind Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87536116-4f1f-4bf0-bdc7-99cd2b88fb1f · inbound
Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Evaluating Large Language Models in Theory of Mind Tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a7d67963-12f6-4b35-8730-ccb355fcadda · inbound
Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Evaluating Large Language Models in Theory of Mind Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa49467c-29a1-4cfd-aefe-e89034e35f02 · inbound
Referential ambiguity and clarification requests: comparing human and LLM behaviour Evaluating Large Language Models in Theory of Mind Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1dbb0e-a190-4d4b-b6bb-4a1d5eb046ee · inbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ae5e3d-4528-42bc-aa53-5d6a2632d481 · inbound
Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task Evaluating Large Language Models in Theory of Mind Tasks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4da330a-b34e-4de7-b732-7b850c1a77d2 · inbound
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Evaluating Large Language Models in Theory of Mind Tasks
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8efacc87-1684-4144-ad46-b7fe1dae4705 · inbound
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology Evaluating Large Language Models in Theory of Mind Tasks
Reference 160
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b3c992-78f1-4148-a060-f5a327d8142f · inbound
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue Evaluating Large Language Models in Theory of Mind Tasks
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac3479e-559f-4eb7-959a-cad129263656 · inbound
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b59a820-a089-44ad-9167-90df738e3c22 · inbound
One Model, Two Minds: A Context-Gated Graph Learner that Recreates Human Biases Evaluating Large Language Models in Theory of Mind Tasks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61906d8-b39e-4735-a074-1cc1184d5f06 · inbound
Designing Psychometric Bias Measures for ChatBots: An Application to Racial Bias Measurement Evaluating Large Language Models in Theory of Mind Tasks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6534f454-410e-42ce-b984-86263344e4fa · inbound
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? Evaluating Large Language Models in Theory of Mind Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa7ea78-7e23-4709-9fdc-7cc218331595 · inbound
Tacit Coordination of Large Language Models Evaluating Large Language Models in Theory of Mind Tasks
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da056b69-54ac-4304-8a9c-54bad4205d8f · inbound
Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web Evaluating Large Language Models in Theory of Mind Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae4278fe-fc73-4da0-bd18-f87ac078a6c0 · inbound
Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Evaluating Large Language Models in Theory of Mind Tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 69956e50-598b-462a-8ec1-4cf13063cfcc · inbound
Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation Evaluating Large Language Models in Theory of Mind Tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27781969-6e61-4d58-98c2-00c8f521e710 · inbound
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1cae9632-b97b-423a-ad37-541a76ead2be · inbound
Don't Make the LLM Read the Graph: Make the Graph Think Evaluating Large Language Models in Theory of Mind Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9840fad6-02c6-48e7-8c51-7f9b00268ce4 · inbound
StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning Evaluating Large Language Models in Theory of Mind Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1629171-dab9-4425-8ed1-1df6d9c01445 · inbound
Truth or Tribe: How In-group Favoritism Prioritize Facts in Persona Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22ca6735-abd6-44bf-8b26-e1bdb4dff4d2 · inbound
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment Evaluating Large Language Models in Theory of Mind Tasks
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35043146-8204-41b9-a520-3ac5905ff0ac · inbound
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50593306-9772-43de-bc61-32be8d8b2cae · inbound
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c123530-e486-449c-967e-0a5232c58e65 · inbound
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code Evaluating Large Language Models in Theory of Mind Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 986fe9de-1b21-47b2-93e9-f682ad720dd5 · inbound
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 966f203d-06b8-4cae-ad0b-5488cb8a1429 · inbound
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence Evaluating Large Language Models in Theory of Mind Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f135b0a1-233c-4de2-863b-acb5a2b5b6aa · inbound
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence Evaluating Large Language Models in Theory of Mind Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e791bbf4-78ef-4352-a83d-62345b5cff33 · inbound
Evaluating Large Language Models in a Complex Hidden Role Game Evaluating Large Language Models in Theory of Mind Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ecab5c85-5f2a-4334-a661-3994de44ea3f · inbound
Voluntary Collusion with Secret Tools in Competing LLM Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b657bf33-1ed2-4e48-8b6b-d119c3727766 · inbound
AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d1685c7-b646-4a5d-881c-45cf02d2f7c9 · inbound
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9c23742-28e3-48dd-a66c-20ddd3e4b0ac · inbound
A Survey of Large Language Models for Perception and Measurement of Human Psychology Evaluating Large Language Models in Theory of Mind Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eb0b92dc-7025-4496-a19e-293b60882bf3 · inbound
Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs) Evaluating Large Language Models in Theory of Mind Tasks
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 19ab65e3-ada9-46b4-8062-35ecadf5645e · inbound
ToxiREX: A Dataset on Toxic REasoning in ConteXt Evaluating Large Language Models in Theory of Mind Tasks
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a06b16e8-868d-4adb-af30-780ec62111cc · inbound
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Evaluating Large Language Models in Theory of Mind Tasks
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c8683c41-d09b-4a34-ba3a-335b740fdc83 · inbound
Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents Evaluating Large Language Models in Theory of Mind Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation da795e3f-c986-46e1-89fa-c232d5ec6918 · inbound
Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework Evaluating Large Language Models in Theory of Mind Tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 84814518-9829-4235-b250-ee316385be28 · inbound
AgentSociety 2: An Integrated Research Environment for Executable Social Science Evaluating Large Language Models in Theory of Mind Tasks
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e8c9ee2-1ef8-43fc-ab00-8086c5a629c2 · inbound
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Evaluating Large Language Models in Theory of Mind Tasks
Reference 164
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b814d5-e1b2-4821-8e66-d3fecded039f · inbound
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Evaluating Large Language Models in Theory of Mind Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.