Pith. sign in

Paper Citation Record · LEDGER

ToMBench: Benchmarking Theory of Mind in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2402.15052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.15052 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:00:46.176692Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.697910Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 096fb5c0-c6ed-41d8-b1f3-d3a05a82b008 · inbound

SocialEval: Evaluating Social Intelligence of Large Language Models cites this paper.

SocialEval: Evaluating Social Intelligence of Large Language Models ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:46.176692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:00:46.176692Z digest=sha256:c8b91a5b54515aeb195dff635c1830b2780da17f016b7f30d67c2c858c990d29

Observation 70729dab-1315-4317-ac4d-b0a6e649ebde · inbound

XToM: Exploring the Multilingual Theory of Mind for Large Language Models cites this paper.

XToM: Exploring the Multilingual Theory of Mind for Large Language Models ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:24.242068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:24.242068Z digest=sha256:cda1adbb81c5b2f674079c69266bda3bbed8680cebb2fe07a72120d74e8cbc72

Observation 6c4fd5b4-864e-404f-9bd9-99e845e896e4 · inbound

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs cites this paper.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:29.167252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:51:29.167252Z digest=sha256:75a886f2b156aca16e9bef464f2b59c2ddc6dfc2a5695af6f582f80e41a5fa0e

Observation 4e7260fd-7c34-4e43-9e8a-116036605cd2 · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.997496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.997496Z digest=sha256:1ff838ac8ad0a76d3284898b73795c4b59a2b3ebf189fa186351293abadf3574

Observation aa3e38bd-9013-4341-a854-eefd00f70df9 · inbound

H2HTalk: Evaluating Large Language Models as Emotional Companion cites this paper.

H2HTalk: Evaluating Large Language Models as Emotional Companion ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:25.488814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:25.488814Z digest=sha256:3a533e24dacb0dafb63d0952517ca4b09aa4c44c31e2ec8eb796bf10a85e3f3d

Observation 225b74d8-5bc3-4b11-9035-eabe25557369 · inbound

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue cites this paper.

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T11:46:53.197894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:46:53.197894Z digest=sha256:fc92134327ed0897fb84f4cc49eba5587ba38d1cf7e8b3adf5697651ce6d888d

Observation de63b124-760f-4767-aaef-880f32043c7c · inbound

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? cites this paper.

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:42:02.985944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:42:02.985944Z digest=sha256:a0cc89f77f3f90f0ef215dba119d202d1fdb0874f4f4b0fa2243fcc4bf3ab0c0

Observation 017fe9d4-f3df-4f66-9bb6-2c19ab37ac0e · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.848790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:ef32606cd8791856e14e41cad2125ca684bae53ddb252c06bbe816c32a5060c7

Observation cd73f55d-325c-43f4-8121-7eaab50b55a5 · inbound

Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation cites this paper.

Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:56:05.796036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:52:22.319870Z digest=sha256:5a79585aa972e897f10333397286de773dc78b550336b46682c621e1dd12c773

Observation 9bc8158f-526b-4759-9fc6-9785713b76ab · inbound

Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues cites this paper.

Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:38.897667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:19:57.347494Z digest=sha256:0e6d65b6d3e1dfe251a8fda2f50c9dcbbd56527b46f2d37eb0c8b412a8792d27

Observation 5ee50c9b-2199-4444-828f-204af964f158 · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.252198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:8c9ed193fd2620c2c228c5c5f5c0d3987f11fa637bf500defdc9cc139ad7f943

Observation e999a52a-b47f-41d9-9018-4f6ad92da217 · inbound

Reinforcing Human Behavior Simulation via Verbal Feedback cites this paper.

Reinforcing Human Behavior Simulation via Verbal Feedback ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:24:02.458142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T07:21:48.649289Z digest=sha256:473d7db047ccd3695a32e58bf7fa215f8d8549601ca1c3c484450b1ada519e6e

Observation a09f3e32-ca1f-419c-bb34-e7844e7e3107 · inbound

Social World Model for Lifelong Social Intelligence cites this paper.

Social World Model for Lifelong Social Intelligence ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:37.152307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:29:11.743627Z digest=sha256:be52181d8225e242a5729d84872387252a85f4ba1cb01eabfe5aaeccb21e627b

Observation a5d78b53-6416-4a3c-bd8b-7996d8f73c68 · inbound

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes cites this paper.

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.699869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:22:34.367119Z digest=sha256:d4df3d8bd8234658ecc1313504e588586e0035479c0f1892cb22560f8e8e38b2

Observation 5c4dded9-82ea-4717-99f6-3212b66dfe39 · inbound

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes cites this paper.

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T10:35:09.399899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:35:09.399899Z digest=sha256:855cff6845982ad8a8148ba346b689e39a64a05fd7320a94dba54d311f566a90

Observation 29f06851-5dc7-4c70-875c-5fbd2e7f7f8c · inbound

Mental World Modeling cites this paper.

Mental World Modeling ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T11:07:38.423077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:07:38.423077Z digest=sha256:8c1e5cb09d2ca5b57a5c0b0f81a2073b0776c837758e13db8616040401f30548

Observation 1d3a5ffb-3472-4559-8c2a-53f637b92169 · inbound

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs cites this paper.

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T08:29:29.096107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:29:29.096107Z digest=sha256:d9dc3051537fed9e7382002448dcccf1b1d79836cf91a8b9e0278056e3c2bf16