Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2302.08399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.08399 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:22:45.056255Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

79
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 74814fe1-819c-4f67-a8b1-7db06893d7b9 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:46:03.367546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:a5044d5b5381eb5e6fe7bd379c0aec6fcebf07896cf32bfa2da9556c6a6a5115

Observation 1649c9ec-82ab-4eb1-a22c-e240d09e33e8 · inbound

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks cites this paper.

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:45.056255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:22:45.056255Z digest=sha256:34473a915e2f092b476d2358a3ce53b69c7489048858035c9599cd92373b0adf

Observation 2338b580-a693-4d3d-a534-5ebe96e9046b · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.382305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:b19550e6c1b965e86ec01cd3ce01943a6781ef48d70b9547b3be677243f454ff

Observation 7db379fd-99a6-4e02-8f69-2e81ad8da768 · inbound

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States cites this paper.

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:29.153547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:29.153547Z digest=sha256:4293b3099776d715fe5b518954694ff51402d6309a5b072e4af1137e6f30ece2

Observation 8277298c-d1ee-4270-87dc-b7bc6b8f79f9 · inbound

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes cites this paper.

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:14.922995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:44:14.922995Z digest=sha256:fbe8d26d243e617aa54c452cdbc314bfc0439af1f1e7ce414c2b06a7e44fa7e1

Observation 1c908b8d-0609-4f01-afd6-24c7428bea7e · inbound

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation cites this paper.

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:25.885019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:25.885019Z digest=sha256:c3bf54db0e56a8e243c41489ef98fb242d7d01c63ed7226f713e7f1be37eb8fc

Observation 54814bd8-0158-4116-890a-620cbc0a365d · inbound

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs cites this paper.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:30.735912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:51:30.735912Z digest=sha256:ecdc687de05f6188fe6f074f916d0bbb51339ba0182b9c32baa221829189a750

Observation 4f45d0ac-3793-4ce4-b70e-57931e169a76 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:26.225894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:26.225894Z digest=sha256:4a24fd17dc503749be06b4d4f9f9ac7ef345ebe9916df180712bd12e778098e6

Observation f6f8376c-1c0d-4adf-a669-f81a9d92728e · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.406955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:17e4289771ca4c9282c2df538be4012584c274534d7729a6e08a02954f4e4e00

Observation 689c85c9-638d-403a-a4bb-8fd9186afa97 · inbound

Mechanistic Interpretability Needs Philosophy cites this paper.

Mechanistic Interpretability Needs Philosophy Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.392344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T23:49:19.683025Z digest=sha256:f49df8d5e0a89e1b1c8d81ffc5d8cf146c464711dd9380acc30e6cbe22201ff7

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:721860dbf493a7a2d5f64abe647c29b053cc5ef5fa8b172de2671340ce7c0d30

Observation afb47a2e-fa82-44b8-89da-0efbba159c7d · inbound

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration cites this paper.

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.124597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:23:56.627581Z digest=sha256:261122302edbe637cef5d79c883c4e30b9919f75a41fbd129718291798d87ed0

Observation 3f912863-fd85-44ee-b125-cddc8c6d7130 · inbound

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning cites this paper.

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:57.810009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:57.810009Z digest=sha256:cf9d4da7bf621377e8b13909b2dc7d71fb63a1ff46130595bec3bc73e29d60b0

Observation 53995b7b-2978-407b-86c7-8db638d4f317 · inbound

Strategy Adaptation in Large Language Model Werewolf Agents cites this paper.

Strategy Adaptation in Large Language Model Werewolf Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:43.283347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:42:43.283347Z digest=sha256:20f5b50c38a05c291fdf6aff4514cc0c394101b15a67cecb7bd883ca0fd59de7

Observation 922a93a1-2cda-4220-8252-7824287ecfff · inbound

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning cites this paper.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:7527302ccb00db3f881f923e0ac8f3cb91cc5ba12b8c93435130da9e679911ce

Observation d998f9ac-10d9-484d-af9f-ab6379ed77fd · inbound

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? cites this paper.

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T05:53:21.310753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:53:21.310753Z digest=sha256:277ccb29a72d7111f9d26f87ce6e08724c0cb1c2e02dc137fc6d584673ef4fbc

Observation c3e10985-b2d6-494f-b893-85e0af453ac7 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:31:29.424187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:c8e99d40b4159c92decf4549c13f705b143175740aacfe2ea490b612d34760ca

Observation 989d248c-1087-4cf8-9957-3f2370b93312 · inbound

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It cites this paper.

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.086520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:33:50.078604Z digest=sha256:61cdfb5fd723f611a44dd1e2a3d309f854dd2ca11a905b84ce9a1353f5068a78

Observation 8a0be8ec-6c49-4634-9eaa-74e7ef352908 · inbound

Network Effects and Agreement Drift in LLM Debates cites this paper.

Network Effects and Agreement Drift in LLM Debates Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:46:08.057855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:50:20.833398Z digest=sha256:38540b659e67a4bcef3309f997301046cab32b1a05aa76696dfbdd070be3b50c

Observation aa691441-60a7-42d2-a723-fa4ff1cfaab6 · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.323113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:581929fd20bda8e54c1cb0883839384a582230db0e2af782c033c0c9c1a2e50d

Observation e6195f40-4131-4d09-bae1-327a0408f14f · inbound

Impact of Task Phrasing on Presumptions in Large Language Models cites this paper.

Impact of Task Phrasing on Presumptions in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:11.361474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:14:04.840446Z digest=sha256:aa10914d4890cd0cb8096a5a4b34bbdbb043878f237530b1f2d2d27da6621d70

Observation a47a5663-6d9e-4e43-934a-5a7cc0bae300 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T02:24:40.718851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T15:36:36.257870Z digest=sha256:ca55a5715ac95033ea3f0fbeceb9e99ebff0eec340592710f60f0d48492b0e09

Observation b1b96d59-26d0-4dcb-8ee1-76d52ab7bcd9 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T18:28:58.421710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:26:14.265380Z digest=sha256:309e488c05a0241a84049b0abc19702200e9c40a7a56e12a5515304b2e331f78

Observation c698d33d-198d-4019-b869-0e9a6f28604c · inbound

ProactBench: Beyond What The User Asked For cites this paper.

ProactBench: Beyond What The User Asked For Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:16:16.052376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T02:14:01.145443Z digest=sha256:874c7ae3fdd21da1a7e7a732e393d14c76cc567b825cbae3600247111ed81d4e

Observation 910cca44-b35e-4301-8c36-71eeaef84d61 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.203788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:59:52.659739Z digest=sha256:82814948f1ed7936f70c80337bbb42d144271c0a671ce620e94675e5ebdd58e3

Observation f2bcdd2b-bc34-427e-bc2c-433fee88eb68 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:29:12.834145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T23:27:45.394730Z digest=sha256:6358ac46f86325e5b362190ae347dee91ed9a6af3fcbe59c047033114558ef6a

Observation 705ce2c7-cab5-4b35-9270-5e97b2bb7a40 · inbound

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue cites this paper.

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:39:03.445907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T21:36:51.291779Z digest=sha256:dcf0b517574801f15991fc2ab352a17eee29d688f638c93f430d66bb5ff8a326

Observation 64bedbf5-ef58-4daf-9b23-796ad86f1c62 · inbound

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance cites this paper.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:32:49.729980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T22:29:17.058647Z digest=sha256:dc22c6d3bb8e01d47afeff428f53144b747bb7946edbd2c0c2b0ba41c05af476

Observation ec8f5611-4a4a-4e04-b445-5b83bcdfc64a · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.895524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:326066c3588d974ca4fcbb3e347b51d62c6383fc7dc2262bee97215e339b0296

Observation 07044555-5fc0-426b-a02a-1543268c205e · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.264689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:7be9df795a32094873220ae8b05807c5158be15498261b4d75c5882981aba31c

Observation c199515d-e77c-4adc-b2b8-d2da8e10ee66 · inbound

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models cites this paper.

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.656455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:41:35.532363Z digest=sha256:b86498fee14fbf083937b95dc7bb39e7588b7e2fdae77e5427b12495e8319397

Observation 33c1074d-40a5-4d69-b625-f1871ef006f2 · inbound

Voluntary Collusion with Secret Tools in Competing LLM Agents cites this paper.

Voluntary Collusion with Secret Tools in Competing LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.780449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T17:01:22.732390Z digest=sha256:3222aaac1ee7bb50717dd14c5c23a45b506ad9de6c97afbcbf7556f581507c7b

Observation 2c7cdacb-ec2a-40eb-a818-b7ef9bc8b574 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.213072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:eada0625e0e1a7bf5c30809fcb2483d0f96c706c8fc589a72e26a3f7a45833ad

Observation 578d6264-66cf-4fb5-87eb-2151f969a46d · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.961583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:eaba40bbad397d1068fc4119227e21e3bbc7a7f443c061a7bbabda98a6ef0a77

Observation b8572d77-fe79-49a8-b0f7-b8f84b5d9bab · inbound

MindZero: Learning Online Mental Reasoning With Zero Annotations cites this paper.

MindZero: Learning Online Mental Reasoning With Zero Annotations Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:10.637583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T21:57:17.194711Z digest=sha256:bc726d30b1a7f25a08c99b72e99e1d66555ef5b2718936b315d33eb1d25c6131

Observation c73e78d3-a925-4449-aed6-0fa229f085fe · inbound

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents cites this paper.

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:56.983506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T02:07:07.135522Z digest=sha256:e193f71e992bc6395a605a8f84440f5400a2184f55e03b5cb15396a28c393bde

Observation 794d8033-5c5d-4ab4-b63f-2df312c469d7 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.207742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:8e78da8fe625852e54aed8b65b1bb8d22bd23e44c545dda65aa076cc485781fa

Observation ad45c732-6d1f-46d4-b57d-af9c6be58ae5 · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.444533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:39:02.190984Z digest=sha256:c8a8604264f7568116f3aec3e77396b11bf0da8dceff9f50055e525f683e89af

Observation c6e2bd2b-00cc-485f-a59f-eebca50ca35b · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T11:47:18.453636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:47:18.453636Z digest=sha256:31b9ae4b2cc4fcfc4556452fb43cbeb3ca8c195c6af8b93f664540530eba7ad0

Observation d0716f81-f13a-455e-8c32-d9352e95038c · inbound

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning cites this paper.

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.702392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:44:31.919126Z digest=sha256:a066332978590999a22b0032ac2e2bd45079c31f1f5183de99035beb63e11b83

Observation 1d834461-315d-4c78-96e1-7941103a4f1c · inbound

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence cites this paper.

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T11:11:46.562648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:11:46.562648Z digest=sha256:11c9cd37388ade288b66cf2dd38c5566149b7590ea6f2117538713ac28664f46

Observation 3e6295fa-ac46-410a-bd95-8f5575fd0c36 · inbound

A Survey of Large Language Models for Perception and Measurement of Human Psychology cites this paper.

A Survey of Large Language Models for Perception and Measurement of Human Psychology Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:04:57.157095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T16:59:25.825681Z digest=sha256:fda32817af41fdb56e7fa451761799fa6be0e20855c4aaabd2f8b8335b1430bd

Observation ec190ba6-5f1c-4c6d-a5a6-a1d2387351bf · inbound

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure cites this paper.

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:45.975167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:10:46.067374Z digest=sha256:ad0d459a3f8e789aa8d0d1bfb887a39ac15e76b0964888c15310fe274073444d

Observation c145cecc-70a6-4ffd-a569-88c561c2dc89 · inbound

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs cites this paper.

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.392659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:48:58.181882Z digest=sha256:0436e857578418afcf365d65ef72bb2a6c04765ee05cae041ddbebf330078d66

Observation 9d1084a6-4b77-4b2a-8fa4-6cff5ffc887a · inbound

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models cites this paper.

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T01:34:08.819629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T01:32:40.504759Z digest=sha256:73c42f1435f37b89dd3ed92e084906f78be8807a8a5e10399c94da9cff5b985e

Observation 83cdd38f-ff5c-4b51-a867-55f095bd3ad4 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.709551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:abcc8e69445648251b8879ceb5d07d87769dfb8505ce8493ad612308a3db50da

Observation 22d2a7af-efa9-4cad-abcc-2ac87a1b027b · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T10:15:59.479435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:15:59.479435Z digest=sha256:a5ce17e54cbd1e30193cf7a04b9776aeff29c68deab66c112370382cc727ed4c

Observation 86c5abea-47e1-4765-9ee2-fadab25216b8 · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:01.579186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:15:01.579186Z digest=sha256:be2965fb4f0bb0fdc4a7fdd8ca735eb98708383c8e3dd0da4f2975424fb90eac

Observation 8276b524-31b1-4e7c-baa1-268296de9bca · inbound

Belief-reality separation lives in routing over a shared value slot in language models cites this paper.

Belief-reality separation lives in routing over a shared value slot in language models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:25:57.402299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:25:57.402299Z digest=sha256:1a20ebe096fc8a5f61851f7b1ca0dacd32b27cbadeed14e5f4c2f06d36396895

Observation 60ace4c0-58ee-40fa-a436-d4e569ff3a88 · inbound

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt cites this paper.

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T02:43:30.604719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:43:30.604719Z digest=sha256:c1c4ec78d9897d664936ed56e6bf49f6273cdbe81706376ca8d6ca764bd3e577

Observation 4429bffe-0e6e-47c8-afce-fa78f5848f81 · inbound

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments cites this paper.

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:46:40.996070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:46:40.996070Z digest=sha256:285c40b8a9745f4683bb519c41430e74c59ceab0cd1f2edf91ec2cfc61880452

Observation b915499d-f861-40a8-a962-dcbf650d71d8 · inbound

Perceived AGI: Believability as Dimensional Completeness, Not Capability cites this paper.

Perceived AGI: Believability as Dimensional Completeness, Not Capability Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:09:15.497907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:09:15.497907Z digest=sha256:0f35e0e424c3dfe41f887a5cd02b3c591e291b13f1dac1437cfaa9992e3e8634

Observation 86e59290-b4d6-4b56-8747-c00352253314 · inbound

Mental World Modeling cites this paper.

Mental World Modeling Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T11:07:38.427392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:07:38.427392Z digest=sha256:e72a13a74bd86dc9810b24d698ad165224be78abddc32828004fb47dec1e2399

Observation efd59943-ef7c-46cf-a890-1f6a849ee043 · inbound

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning cites this paper.

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:58.012382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:58.012382Z digest=sha256:09f6f75eff0f27424efbe523576db3a0921038ebe27f381e07845f03fbbc70c4