Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2302.08399.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:22:45.056255Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
79
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 74814fe1-819c-4f67-a8b1-7db06893d7b9 · inbound
GAIA: a benchmark for General AI Assistants Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1649c9ec-82ab-4eb1-a22c-e240d09e33e8 · inbound
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2338b580-a693-4d3d-a534-5ebe96e9046b · inbound
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7db379fd-99a6-4e02-8f69-2e81ad8da768 · inbound
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8277298c-d1ee-4270-87dc-b7bc6b8f79f9 · inbound
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c908b8d-0609-4f01-afd6-24c7428bea7e · inbound
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54814bd8-0158-4116-890a-620cbc0a365d · inbound
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f45d0ac-3793-4ce4-b70e-57931e169a76 · inbound
From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f8376c-1c0d-4adf-a669-f81a9d92728e · inbound
Bayesian Social Deduction with Graph-Informed Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 689c85c9-638d-403a-a4bb-8fd9186afa97 · inbound
Mechanistic Interpretability Needs Philosophy Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · inbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb47a2e-fa82-44b8-89da-0efbba159c7d · inbound
Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f912863-fd85-44ee-b125-cddc8c6d7130 · inbound
Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53995b7b-2978-407b-86c7-8db638d4f317 · inbound
Strategy Adaptation in Large Language Model Werewolf Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922a93a1-2cda-4220-8252-7824287ecfff · inbound
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d998f9ac-10d9-484d-af9f-ab6379ed77fd · inbound
Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e10985-b2d6-494f-b893-85e0af453ac7 · inbound
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 989d248c-1087-4cf8-9957-3f2370b93312 · inbound
Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a0be8ec-6c49-4634-9eaa-74e7ef352908 · inbound
Network Effects and Agreement Drift in LLM Debates Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa691441-60a7-42d2-a723-fa4ff1cfaab6 · inbound
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6195f40-4131-4d09-bae1-327a0408f14f · inbound
Impact of Task Phrasing on Presumptions in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a47a5663-6d9e-4e43-934a-5a7cc0bae300 · inbound
Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1b96d59-26d0-4dcb-8ee1-76d52ab7bcd9 · inbound
Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c698d33d-198d-4019-b869-0e9a6f28604c · inbound
ProactBench: Beyond What The User Asked For Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 910cca44-b35e-4301-8c36-71eeaef84d61 · inbound
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2bcdd2b-bc34-427e-bc2c-433fee88eb68 · inbound
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 705ce2c7-cab5-4b35-9270-5e97b2bb7a40 · inbound
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64bedbf5-ef58-4daf-9b23-796ad86f1c62 · inbound
Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec8f5611-4a4a-4e04-b445-5b83bcdfc64a · inbound
Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07044555-5fc0-426b-a02a-1543268c205e · inbound
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c199515d-e77c-4adc-b2b8-d2da8e10ee66 · inbound
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33c1074d-40a5-4d69-b625-f1871ef006f2 · inbound
Voluntary Collusion with Secret Tools in Competing LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c7cdacb-ec2a-40eb-a818-b7ef9bc8b574 · inbound
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 578d6264-66cf-4fb5-87eb-2151f969a46d · inbound
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8572d77-fe79-49a8-b0f7-b8f84b5d9bab · inbound
MindZero: Learning Online Mental Reasoning With Zero Annotations Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c73e78d3-a925-4449-aed6-0fa229f085fe · inbound
AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 794d8033-5c5d-4ab4-b63f-2df312c469d7 · inbound
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad45c732-6d1f-46d4-b57d-af9c6be58ae5 · inbound
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6e2bd2b-00cc-485f-a59f-eebca50ca35b · inbound
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0716f81-f13a-455e-8c32-d9352e95038c · inbound
Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d834461-315d-4c78-96e1-7941103a4f1c · inbound
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e6295fa-ac46-410a-bd95-8f5575fd0c36 · inbound
A Survey of Large Language Models for Perception and Measurement of Human Psychology Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec190ba6-5f1c-4c6d-a5a6-a1d2387351bf · inbound
When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c145cecc-70a6-4ffd-a569-88c561c2dc89 · inbound
Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d1084a6-4b77-4b2a-8fa4-6cff5ffc887a · inbound
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83cdd38f-ff5c-4b51-a867-55f095bd3ad4 · inbound
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22d2a7af-efa9-4cad-abcc-2ac87a1b027b · inbound
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c5abea-47e1-4765-9ee2-fadab25216b8 · inbound
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8276b524-31b1-4e7c-baa1-268296de9bca · inbound
Belief-reality separation lives in routing over a shared value slot in language models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ace4c0-58ee-40fa-a436-d4e569ff3a88 · inbound
The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4429bffe-0e6e-47c8-afce-fa78f5848f81 · inbound
Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b915499d-f861-40a8-a962-dcbf650d71d8 · inbound
Perceived AGI: Believability as Dimensional Completeness, Not Capability Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e59290-b4d6-4b56-8747-c00352253314 · inbound
Mental World Modeling Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd59943-ef7c-46cf-a890-1f6a849ee043 · inbound
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.