Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2302.08399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.08399 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:04:09.205936Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

79
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 74814fe1-819c-4f67-a8b1-7db06893d7b9 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:46:03.367546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:7eedf59ab67560e5e804b0700910c9668e384a5ec67dcad08a093983d69dd1d8

Observation 1c14abec-f2a6-47e7-9146-01f2fa7d51f8 · inbound

Why human-AI relationships need socioaffective alignment cites this paper.

Why human-AI relationships need socioaffective alignment Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-09T11:53:43.845529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:53:43.845529Z digest=sha256:b8d8d439d54dd8a0cd4781d4eacf21bdb9d359933a7d788774e3e8b85994803e

Observation 1649c9ec-82ab-4eb1-a22c-e240d09e33e8 · inbound

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks cites this paper.

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:45.056255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:22:45.056255Z digest=sha256:90c7069adc364100d6e4e0b0f952864bee84c031103055b7383de50af54b79b9

Observation fe4635fa-3599-4b14-a8ce-7833fd8713ad · inbound

Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice cites this paper.

Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:04:09.205936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:04:09.205936Z digest=sha256:631c87e2482190a7ee51b4f79396676c61d3d6e75ba4cdb534bdf797c287630c

Observation 2338b580-a693-4d3d-a534-5ebe96e9046b · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.382305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:d7ca3ef61634da157ad62fe1ce2a5be3ca7cbc52943d74a91ffed83fe3311916

Observation 7db379fd-99a6-4e02-8f69-2e81ad8da768 · inbound

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States cites this paper.

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:29.153547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:29.153547Z digest=sha256:8d9e7c6571d131f7f0630be952465a81e09f07c97dce6f2fbc3513cce5849ca7

Observation 8277298c-d1ee-4270-87dc-b7bc6b8f79f9 · inbound

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes cites this paper.

Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:14.922995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:44:14.922995Z digest=sha256:fa8207bbd4cf6c7934076214c8cce3817643e9c19aad6336dd16b370f24ac359

Observation 1c908b8d-0609-4f01-afd6-24c7428bea7e · inbound

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation cites this paper.

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:25.885019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:25.885019Z digest=sha256:908a1e91ed4ca704e95aa5c5630ae62527b1634245f3c0d4a96ef246c9333676

Observation 54814bd8-0158-4116-890a-620cbc0a365d · inbound

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs cites this paper.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:30.735912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:51:30.735912Z digest=sha256:26067740e8d0f48c3994e16381b70e0f59721e7ced834184d363768b12a7d873

Observation 4f45d0ac-3793-4ce4-b70e-57931e169a76 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:26.225894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:26.225894Z digest=sha256:2295147267dafdd3d18410d6c3a6c6b62c3827d5518188bc2efee704d2ecd36b

Observation f6f8376c-1c0d-4adf-a669-f81a9d92728e · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.406955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:4f75d9145c9accf6e8e22a0779a52fcf353338305d0397538cc40fd76ef834a9

Observation 689c85c9-638d-403a-a4bb-8fd9186afa97 · inbound

Mechanistic Interpretability Needs Philosophy cites this paper.

Mechanistic Interpretability Needs Philosophy Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.392344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T23:49:19.683025Z digest=sha256:638d79007dfb588f85cffc05cdf171ca1a07514b7aac0094b41942b9dc8233a5

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:b664f317cc53cc4b77aba91259ec60d926f7c9ecad452c1ffeba1056f74b24d9

Observation afb47a2e-fa82-44b8-89da-0efbba159c7d · inbound

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration cites this paper.

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.124597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:23:56.627581Z digest=sha256:440819d483a7bfe4a316d08119275172b2763372d12b99d700855a24920bee9b

Observation 3f912863-fd85-44ee-b125-cddc8c6d7130 · inbound

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning cites this paper.

Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:57.810009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:57.810009Z digest=sha256:14ac9f4125e47d51273f066e0b91071c9d44fbf2058bf71d2c6d11aecc3ef210

Observation 53995b7b-2978-407b-86c7-8db638d4f317 · inbound

Strategy Adaptation in Large Language Model Werewolf Agents cites this paper.

Strategy Adaptation in Large Language Model Werewolf Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:43.283347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:42:43.283347Z digest=sha256:4b8fd5b4981314862309d8f949920c42baf7cafcfc061e33a91f318274d5fcd9

Observation 922a93a1-2cda-4220-8252-7824287ecfff · inbound

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning cites this paper.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:bc82aa218e1c10e062559a5c718e3ed55178680be53012bc4e3a23c9d595cf11

Observation d998f9ac-10d9-484d-af9f-ab6379ed77fd · inbound

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? cites this paper.

Memorization $\neq$ Understanding: Do Large Language Models Have the Ability of Scenario Cognition? Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T05:53:21.310753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:53:21.310753Z digest=sha256:34a981f5969d9b807330d91cdbbad06435bcbaeb82ca637a426008bd892c3586

Observation c3e10985-b2d6-494f-b893-85e0af453ac7 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:31:29.424187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:bd5d05bba269160bf712efbeb3074b51397e43634927e6e186e730bf154d0d6e

Observation 989d248c-1087-4cf8-9957-3f2370b93312 · inbound

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It cites this paper.

Gradual Cognitive Externalization: From Modeling Cognition to Constituting It Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:48.086520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:33:50.078604Z digest=sha256:7615c775c09e77e6c687bcc2064b8f611b86bc1630336de455e3fee34d9ac070

Observation 8a0be8ec-6c49-4634-9eaa-74e7ef352908 · inbound

Network Effects and Agreement Drift in LLM Debates cites this paper.

Network Effects and Agreement Drift in LLM Debates Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:46:08.057855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:50:20.833398Z digest=sha256:080aaca31c60e0014865720436ee92808ddd6756f68b00f4618afbebfcaa792b

Observation aa691441-60a7-42d2-a723-fa4ff1cfaab6 · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.323113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:3ec2ce2f07fdf7b806d61a249561df7f62c5e25236dcfc4badcbf6023000c034

Observation e6195f40-4131-4d09-bae1-327a0408f14f · inbound

Impact of Task Phrasing on Presumptions in Large Language Models cites this paper.

Impact of Task Phrasing on Presumptions in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:11.361474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:14:04.840446Z digest=sha256:a11f26c5b216346069bfc8b384f8a3a12106e34f05669e9b4a05f110348c47cc

Observation a47a5663-6d9e-4e43-934a-5a7cc0bae300 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T02:24:40.718851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T15:36:36.257870Z digest=sha256:8ec4531fe21a7c69683eb678d0daef0fad059f7e9c2cf56c56edad9d50369163

Observation b1b96d59-26d0-4dcb-8ee1-76d52ab7bcd9 · inbound

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior cites this paper.

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T18:28:58.421710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:26:14.265380Z digest=sha256:0d8f47a5a7c5c937e3d0b8e73aaf3765433563fea0c553e1e26763018b2bca11

Observation c698d33d-198d-4019-b869-0e9a6f28604c · inbound

ProactBench: Beyond What The User Asked For cites this paper.

ProactBench: Beyond What The User Asked For Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:16:16.052376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:14:01.145443Z digest=sha256:fcd15464500aff246b4477ffaafe1ea233b641cf340b00cc88fa87a28d805bab

Observation 910cca44-b35e-4301-8c36-71eeaef84d61 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.203788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:59:52.659739Z digest=sha256:9576b093266097094067ba97774e962059cf23a98e8f124c24e83bf814ac280d

Observation f2bcdd2b-bc34-427e-bc2c-433fee88eb68 · inbound

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents cites this paper.

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:29:12.834145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T23:27:45.394730Z digest=sha256:ed362566606bc0ae7eaf68f0d7f03890a5529625a9c630afb9b6bcfdeaf06e79

Observation 705ce2c7-cab5-4b35-9270-5e97b2bb7a40 · inbound

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue cites this paper.

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T21:39:03.445907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T21:36:51.291779Z digest=sha256:a7f788e3f2ef7de56068919ef75eea358e6f20d5373bc8251ea55c9bb0ad5c10

Observation 64bedbf5-ef58-4daf-9b23-796ad86f1c62 · inbound

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance cites this paper.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:32:49.729980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T22:29:17.058647Z digest=sha256:ecfc0fd642e67f249e090cb16bee20c6a5f1a308df7f34f26a7de027378e5c66

Observation ec8f5611-4a4a-4e04-b445-5b83bcdfc64a · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.895524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:43f5904cf69bf2f33eea016b780a6e2c86671eca7694a210c59458d507f227ba

Observation 07044555-5fc0-426b-a02a-1543268c205e · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.264689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:affaf718a39e97d11cb62f6b7537a7961699a666efbe6171e5eafbfc860683d7

Observation c199515d-e77c-4adc-b2b8-d2da8e10ee66 · inbound

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models cites this paper.

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.656455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T04:41:35.532363Z digest=sha256:6c4a259e9d59de10711adc98f1dae620313b7c910d07387fac780ddc49cd2f0d

Observation 33c1074d-40a5-4d69-b625-f1871ef006f2 · inbound

Voluntary Collusion with Secret Tools in Competing LLM Agents cites this paper.

Voluntary Collusion with Secret Tools in Competing LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.780449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T17:01:22.732390Z digest=sha256:0ab7eeec42316cd13627894d085aa57e9670e28816b863e7f98152ca31626510

Observation 2c7cdacb-ec2a-40eb-a818-b7ef9bc8b574 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.213072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:d95ea4c7c58bcb89dd8daf40f08b862a4e0e7389d38b74d5ba09ef7521650c75

Observation 578d6264-66cf-4fb5-87eb-2151f969a46d · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.961583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:817cb82b7e5f6d139acf2bba3a452f5ca5a71a4674017253b7cdbdd174a34417

Observation b8572d77-fe79-49a8-b0f7-b8f84b5d9bab · inbound

MindZero: Learning Online Mental Reasoning With Zero Annotations cites this paper.

MindZero: Learning Online Mental Reasoning With Zero Annotations Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:10.637583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T21:57:17.194711Z digest=sha256:afe80b5d82fa3545c8608a77febf9f58b391739d19ab475fd0039b85240e82e6

Observation c73e78d3-a925-4449-aed6-0fa229f085fe · inbound

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents cites this paper.

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:56.983506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:07:07.135522Z digest=sha256:cb1dbd1f76e772b71d2c5bfc3f35974a8dd73c853fad2363bbc5ab43fed48756

Observation 794d8033-5c5d-4ab4-b63f-2df312c469d7 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.207742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:c8126c4cec21425f2f0a8a435fdbf15b2d1c22da6baf7f374d9878815203bdaa

Observation ad45c732-6d1f-46d4-b57d-af9c6be58ae5 · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.444533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:39:02.190984Z digest=sha256:4a5ca8c03abff5c8c675ec4b25c16a2a2c550a8544d3f56e9e1207747b700516

Observation c6e2bd2b-00cc-485f-a59f-eebca50ca35b · inbound

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism cites this paper.

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T11:47:18.453636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:47:18.453636Z digest=sha256:f6425a1fc38168a9c7d794e671e3fc3489f7f136887753f6346856822c3baaf7

Observation d0716f81-f13a-455e-8c32-d9352e95038c · inbound

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning cites this paper.

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.702392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T06:44:31.919126Z digest=sha256:6bda52000edd169c1b97fe8089d87af9480a3b34f5ee0f812a90df5b9cee8bad

Observation 1d834461-315d-4c78-96e1-7941103a4f1c · inbound

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence cites this paper.

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T11:11:46.562648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:11:46.562648Z digest=sha256:761e0b94d864861d3623477364865c2cd89bfb811e73761e5709871a269c5b1a

Observation 3e6295fa-ac46-410a-bd95-8f5575fd0c36 · inbound

A Survey of Large Language Models for Perception and Measurement of Human Psychology cites this paper.

A Survey of Large Language Models for Perception and Measurement of Human Psychology Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:04:57.157095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T16:59:25.825681Z digest=sha256:dd6da76b687f8b9a4d6a0a5b455a8183f329dfadb6ae14e47c66458dbeb11e1b

Observation ec190ba6-5f1c-4c6d-a5a6-a1d2387351bf · inbound

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure cites this paper.

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:45.975167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:10:46.067374Z digest=sha256:af0cc4a14d6bf6d4b6951768086da1089dd952d11b8cf6dc294fe9d70005a64b

Observation c145cecc-70a6-4ffd-a569-88c561c2dc89 · inbound

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs cites this paper.

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.392659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T04:48:58.181882Z digest=sha256:346f85dd3ab7c3b7d6dce6ce16566912b8eab202d3fe250aafd188d67c4fb045

Observation 9d1084a6-4b77-4b2a-8fa4-6cff5ffc887a · inbound

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models cites this paper.

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T01:34:08.819629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T01:32:40.504759Z digest=sha256:17ea56558ba5b14d2aacc43b161155aa971e267fd07a124e4455db80a6676e85

Observation 83cdd38f-ff5c-4b51-a867-55f095bd3ad4 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.709551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:0ba9c409daa94e1a643971e38b33a896dc7f641ff0a03238defbc33e192cdbd8

Observation 22d2a7af-efa9-4cad-abcc-2ac87a1b027b · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T10:15:59.479435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:15:59.479435Z digest=sha256:bb423ee6be0456e8b285ff5cee7e91d356a3e7b87164101f2880e37b6961a82e

Observation 86c5abea-47e1-4765-9ee2-fadab25216b8 · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:01.579186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:15:01.579186Z digest=sha256:84530cbb0a00ab273578e60a298de65f098d14629009b4dffc02974a52612538

Observation 8276b524-31b1-4e7c-baa1-268296de9bca · inbound

Belief-reality separation lives in routing over a shared value slot in language models cites this paper.

Belief-reality separation lives in routing over a shared value slot in language models Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:25:57.402299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:25:57.402299Z digest=sha256:e532f24508763608fba441ebf3c239344cd705f580d1429061dfe355e17c28c9

Observation 60ace4c0-58ee-40fa-a436-d4e569ff3a88 · inbound

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt cites this paper.

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T02:43:30.604719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:43:30.604719Z digest=sha256:ad1aaa90f9218c18a0a6b399a53f8651bb01773e691a66af1d9064d5d160d200

Observation 4429bffe-0e6e-47c8-afce-fa78f5848f81 · inbound

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments cites this paper.

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:46:40.996070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:46:40.996070Z digest=sha256:cf688006d8ce5b65d07745233888845b6223de5c0b1ba404368cbc377dbefb72

Observation b915499d-f861-40a8-a962-dcbf650d71d8 · inbound

Perceived AGI: Believability as Dimensional Completeness, Not Capability cites this paper.

Perceived AGI: Believability as Dimensional Completeness, Not Capability Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:09:15.497907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:09:15.497907Z digest=sha256:3157f0475cd0fcca3aa08ca12f904169b4bcff03ad29075da64aabdd0a27f858

Observation 86e59290-b4d6-4b56-8747-c00352253314 · inbound

Mental World Modeling cites this paper.

Mental World Modeling Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T11:07:38.427392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:07:38.427392Z digest=sha256:04ce439be46b1688ed4c8addcb6afa60500a4e0cc2bf98a6451b71fc373eafdc

Observation efd59943-ef7c-46cf-a890-1f6a849ee043 · inbound

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning cites this paper.

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:58.012382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:57:58.012382Z digest=sha256:8ccd1650173f95fdc6b846ffa726d2dfa24109cac009110401b412313d2e30c1