Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2302.08399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.08399 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:53:43.845529Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

79
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 74814fe1-819c-4f67-a8b1-7db06893d7b9 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:46:03.367546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:79f943ddd8cb766840affc277953ae79f1f230a41e8659218f2636e74b4cda6b

Observation 1c14abec-f2a6-47e7-9146-01f2fa7d51f8 · inbound

Why human-AI relationships need socioaffective alignment cites this paper.

Why human-AI relationships need socioaffective alignment Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-09T11:53:43.845529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:53:43.845529Z digest=sha256:b8d8d439d54dd8a0cd4781d4eacf21bdb9d359933a7d788774e3e8b85994803e

Observation 1649c9ec-82ab-4eb1-a22c-e240d09e33e8 · inbound

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks cites this paper.

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:45.056255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:22:45.056255Z digest=sha256:471ee5c7df56a1ab7129c67207ffa3d93dcf65b552a351d65ff81b82896050ae

Observation 2338b580-a693-4d3d-a534-5ebe96e9046b · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.382305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:ca7d5360647e8db98a26e2f9fe40ba0861e322b4e2832f88cbb50bece554d1bd

Observation 7db379fd-99a6-4e02-8f69-2e81ad8da768 · inbound

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States cites this paper.

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:29.153547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:29.153547Z digest=sha256:8d9e7c6571d131f7f0630be952465a81e09f07c97dce6f2fbc3513cce5849ca7

Observation 8277298c-d1ee-4270-87dc-b7bc6b8f79f9 · inbound

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes cites this paper.

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:14.922995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:44:14.922995Z digest=sha256:fa8207bbd4cf6c7934076214c8cce3817643e9c19aad6336dd16b370f24ac359

Observation 1c908b8d-0609-4f01-afd6-24c7428bea7e · inbound

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation cites this paper.

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:25.885019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:25.885019Z digest=sha256:908a1e91ed4ca704e95aa5c5630ae62527b1634245f3c0d4a96ef246c9333676

Observation 54814bd8-0158-4116-890a-620cbc0a365d · inbound

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs cites this paper.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:30.735912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:51:30.735912Z digest=sha256:3500cb0e15f3d3c1cec5d8bb002b5de89992630156d2511994e34d8c1c029f09

Observation 4f45d0ac-3793-4ce4-b70e-57931e169a76 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:26.225894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:26.225894Z digest=sha256:2295147267dafdd3d18410d6c3a6c6b62c3827d5518188bc2efee704d2ecd36b

Observation f6f8376c-1c0d-4adf-a669-f81a9d92728e · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.406955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:abc43d48f832b59f0671179a7698cbec554e5ef737ef4344136cec10d3c91e30

Observation 689c85c9-638d-403a-a4bb-8fd9186afa97 · inbound

Mechanistic Interpretability Needs Philosophy cites this paper.

Mechanistic Interpretability Needs Philosophy Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.392344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T23:49:19.683025Z digest=sha256:ae0d1ccfd1b47d8ba98d53f154134db0539d1ea89fdb809fc8bc109f2d0667a3

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:b664f317cc53cc4b77aba91259ec60d926f7c9ecad452c1ffeba1056f74b24d9

Observation afb47a2e-fa82-44b8-89da-0efbba159c7d · inbound

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration cites this paper.

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.124597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:23:56.627581Z digest=sha256:08a57aa7dabfb503e98457896c7602c50f502d8b0acfa9ec4fd521e3f6fc4f2c

Observation 3f912863-fd85-44ee-b125-cddc8c6d7130 · inbound

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning cites this paper.

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:57.810009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:57.810009Z digest=sha256:14ac9f4125e47d51273f066e0b91071c9d44fbf2058bf71d2c6d11aecc3ef210

Observation 53995b7b-2978-407b-86c7-8db638d4f317 · inbound

Strategy Adaptation in Large Language Model Werewolf Agents cites this paper.

Strategy Adaptation in Large Language Model Werewolf Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:43.283347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:42:43.283347Z digest=sha256:ef46671896130d27006fd00cc73a01a73f7c4a76d09648bf7609db4cec90e9c3

Observation 922a93a1-2cda-4220-8252-7824287ecfff · inbound

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning cites this paper.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:bc82aa218e1c10e062559a5c718e3ed55178680be53012bc4e3a23c9d595cf11

Observation d998f9ac-10d9-484d-af9f-ab6379ed77fd · inbound

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? cites this paper.

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T05:53:21.310753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:53:21.310753Z digest=sha256:91cc436eb8ba0999d397fe72ed25bacd389eec0fc77a2ba0acca0607fba1fb10

Observation c3e10985-b2d6-494f-b893-85e0af453ac7 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:31:29.424187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:d0fccc38d15b14d7d59e78c7be1cb466e6c1351c4bbbb81617405bdaa7fb01b3

Observation 989d248c-1087-4cf8-9957-3f2370b93312 · inbound

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It cites this paper.

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.086520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:33:50.078604Z digest=sha256:72dc7292222f515a51141256e468259cc451264c8e62b023c7f51768a3206e2c

Observation 8a0be8ec-6c49-4634-9eaa-74e7ef352908 · inbound

Network Effects and Agreement Drift in LLM Debates cites this paper.

Network Effects and Agreement Drift in LLM Debates Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:46:08.057855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:50:20.833398Z digest=sha256:c1487e89a137738eba1c3f46e7219d2d01f8a069269a336e1e170fb32c97be0f

Observation aa691441-60a7-42d2-a723-fa4ff1cfaab6 · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.323113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:562fe790b4de173ceb3cdc51ba49ffa2bf43c03e41f577c79dff38202b3d6729

Observation e6195f40-4131-4d09-bae1-327a0408f14f · inbound

Impact of Task Phrasing on Presumptions in Large Language Models cites this paper.

Impact of Task Phrasing on Presumptions in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:11.361474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T19:14:04.840446Z digest=sha256:44a8017a440c794d18029707f0b0e0a9432a70694afa4e46fd3c88e0b95e78c4

Observation a47a5663-6d9e-4e43-934a-5a7cc0bae300 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T02:24:40.718851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T15:36:36.257870Z digest=sha256:e6b2d9c4d4b4304f06bdf5e00bb376cd18aa5af6fd3254ecd56faeda1b082de9

Observation b1b96d59-26d0-4dcb-8ee1-76d52ab7bcd9 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T18:28:58.421710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:26:14.265380Z digest=sha256:791f7d3dc260dc05a9ab3a92272efb4b690e55da48c309000ad9b5f3e5e0048c

Observation c698d33d-198d-4019-b869-0e9a6f28604c · inbound

ProactBench: Beyond What The User Asked For cites this paper.

ProactBench: Beyond What The User Asked For Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:16:16.052376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T02:14:01.145443Z digest=sha256:0d276f420b9de849d3b007a098d04b71361a060e81f89432aedc1768cff15587

Observation 910cca44-b35e-4301-8c36-71eeaef84d61 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.203788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:59:52.659739Z digest=sha256:fb159b488ccc88d70435f6fcf1811746f3b3d9d76c1578d33ff5b5060a5e5a4a

Observation f2bcdd2b-bc34-427e-bc2c-433fee88eb68 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:29:12.834145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:27:45.394730Z digest=sha256:610f31a70cbc22ab20df58d97c4ccfde70f330824bd0b89445ab542baa18e10e

Observation 705ce2c7-cab5-4b35-9270-5e97b2bb7a40 · inbound

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue cites this paper.

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:39:03.445907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:36:51.291779Z digest=sha256:ae7aa438c9647512493e76048d1241dfad90a92afe5506747a2eee2204e42c69

Observation 64bedbf5-ef58-4daf-9b23-796ad86f1c62 · inbound

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance cites this paper.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:32:49.729980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T22:29:17.058647Z digest=sha256:c4d528cdc8c4ed719a93b81a80898440ed8cf4ed04b00af8d0de71c3e37973a3

Observation ec8f5611-4a4a-4e04-b445-5b83bcdfc64a · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.895524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:c8233115f25b6151a3866b02fed27ccf682c98fd13002d9903cd419bb7d16b29

Observation 07044555-5fc0-426b-a02a-1543268c205e · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.264689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:ef82e98056c3066d62851c953bf58775cb11b5b5159d6315730c222c657faeb2

Observation c199515d-e77c-4adc-b2b8-d2da8e10ee66 · inbound

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models cites this paper.

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.656455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:41:35.532363Z digest=sha256:dd15974aff37a5d5a23cdb6d2c7ca0f865894fa295cdef07e4636ab23cfd6ddc

Observation 33c1074d-40a5-4d69-b625-f1871ef006f2 · inbound

Voluntary Collusion with Secret Tools in Competing LLM Agents cites this paper.

Voluntary Collusion with Secret Tools in Competing LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.780449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:01:22.732390Z digest=sha256:d7a43351e54ddbb3a5ddd8048cb8049eb6a81254c9bc10cb28e2a5a604fdefc8

Observation 2c7cdacb-ec2a-40eb-a818-b7ef9bc8b574 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.213072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:6acb50172a2c7b83af5410d4353bc59a5116ee2b3f07001660948e1f86446dcb

Observation 578d6264-66cf-4fb5-87eb-2151f969a46d · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.961583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:531ce6236924fc63bcdd6c3ea87b8ec33af4e99b1ae41a894d222a09235898bb

Observation b8572d77-fe79-49a8-b0f7-b8f84b5d9bab · inbound

MindZero: Learning Online Mental Reasoning With Zero Annotations cites this paper.

MindZero: Learning Online Mental Reasoning With Zero Annotations Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:10.637583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T21:57:17.194711Z digest=sha256:76db3b7d1c5d5227d8f1323f3b191bc3c6906f4e0a340540eebcc6f4f13e286b

Observation c73e78d3-a925-4449-aed6-0fa229f085fe · inbound

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents cites this paper.

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:56.983506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:07:07.135522Z digest=sha256:2a0838dd5c6583437bca17d734221f8b35e0a254a668af9b768f09eb01aba0b1

Observation 794d8033-5c5d-4ab4-b63f-2df312c469d7 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.207742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:128b0ab4a8d2eeb78d0a4f56d7c97b81fc4c8bc2e908bb406c8985e851abc5d0

Observation ad45c732-6d1f-46d4-b57d-af9c6be58ae5 · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.444533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:39:02.190984Z digest=sha256:ffec0f4aa28b558a5896fc20c9c14289220bac00de76987893bd9fc35daf2e53

Observation c6e2bd2b-00cc-485f-a59f-eebca50ca35b · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T11:47:18.453636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:47:18.453636Z digest=sha256:f6425a1fc38168a9c7d794e671e3fc3489f7f136887753f6346856822c3baaf7

Observation d0716f81-f13a-455e-8c32-d9352e95038c · inbound

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning cites this paper.

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.702392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:44:31.919126Z digest=sha256:4c5c9fff299cbdfa8c2b30798eb2b7ef36dd4292106c8e094f47d3f7c797a547

Observation 1d834461-315d-4c78-96e1-7941103a4f1c · inbound

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence cites this paper.

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T11:11:46.562648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:11:46.562648Z digest=sha256:761e0b94d864861d3623477364865c2cd89bfb811e73761e5709871a269c5b1a

Observation 3e6295fa-ac46-410a-bd95-8f5575fd0c36 · inbound

A Survey of Large Language Models for Perception and Measurement of Human Psychology cites this paper.

A Survey of Large Language Models for Perception and Measurement of Human Psychology Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:04:57.157095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:59:25.825681Z digest=sha256:57f5b6582447209c80d4e4ea9ab3286db9ba03b06208445b0eed8a502cac41f7

Observation ec190ba6-5f1c-4c6d-a5a6-a1d2387351bf · inbound

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure cites this paper.

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:45.975167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:10:46.067374Z digest=sha256:c7cc67ef926624b952b13a1569477b334d6c96a71233e6e2b774532d1b80d766

Observation c145cecc-70a6-4ffd-a569-88c561c2dc89 · inbound

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs cites this paper.

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.392659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:48:58.181882Z digest=sha256:2d4b6e568d87274b71cfa16899cdd5b7fbe79a6d8fcd9c5e45d0a5c8808b1050

Observation 9d1084a6-4b77-4b2a-8fa4-6cff5ffc887a · inbound

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models cites this paper.

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T01:34:08.819629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T01:32:40.504759Z digest=sha256:39368059d5d9d002cafab46148e8658c07568e9e626ed3c865f02b37bdfa9cf4

Observation 83cdd38f-ff5c-4b51-a867-55f095bd3ad4 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.709551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:b9bf8b94fc06c5bf06530efecfb9b985c96775b5cfb55ea6ae9395b68a7f8ff5

Observation 22d2a7af-efa9-4cad-abcc-2ac87a1b027b · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T10:15:59.479435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:15:59.479435Z digest=sha256:bb423ee6be0456e8b285ff5cee7e91d356a3e7b87164101f2880e37b6961a82e

Observation 86c5abea-47e1-4765-9ee2-fadab25216b8 · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:01.579186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:15:01.579186Z digest=sha256:84530cbb0a00ab273578e60a298de65f098d14629009b4dffc02974a52612538

Observation 8276b524-31b1-4e7c-baa1-268296de9bca · inbound

Belief-reality separation lives in routing over a shared value slot in language models cites this paper.

Belief-reality separation lives in routing over a shared value slot in language models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:25:57.402299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:25:57.402299Z digest=sha256:e532f24508763608fba441ebf3c239344cd705f580d1429061dfe355e17c28c9

Observation 60ace4c0-58ee-40fa-a436-d4e569ff3a88 · inbound

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt cites this paper.

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T02:43:30.604719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:43:30.604719Z digest=sha256:ad1aaa90f9218c18a0a6b399a53f8651bb01773e691a66af1d9064d5d160d200

Observation 4429bffe-0e6e-47c8-afce-fa78f5848f81 · inbound

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments cites this paper.

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:46:40.996070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:46:40.996070Z digest=sha256:3cd02103f0e36a787b8232fbf9387c2af1d4e0ba519e49f2e9d2a4295b1b53d4

Observation b915499d-f861-40a8-a962-dcbf650d71d8 · inbound

Perceived AGI: Believability as Dimensional Completeness, Not Capability cites this paper.

Perceived AGI: Believability as Dimensional Completeness, Not Capability Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:09:15.497907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:09:15.497907Z digest=sha256:3157f0475cd0fcca3aa08ca12f904169b4bcff03ad29075da64aabdd0a27f858

Observation 86e59290-b4d6-4b56-8747-c00352253314 · inbound

Mental World Modeling cites this paper.

Mental World Modeling Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T11:07:38.427392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:07:38.427392Z digest=sha256:04ce439be46b1688ed4c8addcb6afa60500a4e0cc2bf98a6451b71fc373eafdc

Observation efd59943-ef7c-46cf-a890-1f6a849ee043 · inbound

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning cites this paper.

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:58.012382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:58.012382Z digest=sha256:8ccd1650173f95fdc6b846ffa726d2dfa24109cac009110401b412313d2e30c1