Pith. sign in

Paper Citation Record · LEDGER

ToMBench: Benchmarking Theory of Mind in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2402.15052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.15052 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:00:46.176692Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.697910Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 096fb5c0-c6ed-41d8-b1f3-d3a05a82b008 · inbound

SocialEval: Evaluating Social Intelligence of Large Language Models cites this paper.

SocialEval: Evaluating Social Intelligence of Large Language Models ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:46.176692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:00:46.176692Z digest=sha256:6628cc4ce3243daafa63712917a4f82ccbf57902de2765f50b99c1170a49542c

Observation 70729dab-1315-4317-ac4d-b0a6e649ebde · inbound

XToM: Exploring the Multilingual Theory of Mind for Large Language Models cites this paper.

XToM: Exploring the Multilingual Theory of Mind for Large Language Models ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:24.242068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:24.242068Z digest=sha256:4af33f206848884000fae7401ca65d5746372b75abc116a3bdc871903dfc57cc

Observation 6c4fd5b4-864e-404f-9bd9-99e845e896e4 · inbound

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs cites this paper.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:29.167252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:51:29.167252Z digest=sha256:1b126565b1bda8335c05a66601c246b9f98837d62f04a4bb48aa4bf87bb937a0

Observation 4e7260fd-7c34-4e43-9e8a-116036605cd2 · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.997496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.997496Z digest=sha256:5b8d0313015f86f414aedb48cd7da052c75c2cccc3129026059241da586d9d90

Observation aa3e38bd-9013-4341-a854-eefd00f70df9 · inbound

H2HTalk: Evaluating Large Language Models as Emotional Companion cites this paper.

H2HTalk: Evaluating Large Language Models as Emotional Companion ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:25.488814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:25.488814Z digest=sha256:633e75f4804be777e1a19a211b1d3de8e94f7324994ac6219bffa5f9b478f22b

Observation 225b74d8-5bc3-4b11-9035-eabe25557369 · inbound

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue cites this paper.

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T11:46:53.197894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:46:53.197894Z digest=sha256:e6ce0ab1c94f8a3ba777bda37ca6cc2c7cdbfa1e979a90e55239066dd453939a

Observation de63b124-760f-4767-aaef-880f32043c7c · inbound

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? cites this paper.

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About? ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:42:02.985944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:42:02.985944Z digest=sha256:6e4711712a7eb764551c738a7b8093f06c911994bb7532d0c372cb907cbfa818

Observation 017fe9d4-f3df-4f66-9bb6-2c19ab37ac0e · inbound

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum cites this paper.

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.848790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:08:54.560648Z digest=sha256:bdd6db5624754abbf8a3e1cb1152380b8e4bd92d65d9f9d99a24fcde69a42286

Observation cd73f55d-325c-43f4-8121-7eaab50b55a5 · inbound

Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation cites this paper.

Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:56:05.796036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:52:22.319870Z digest=sha256:2f42f9b0a13b5dc65f8fe7ddd15f2b7a7f2f613eae9465b50821d9b71c37f954

Observation 9bc8158f-526b-4759-9fc6-9785713b76ab · inbound

Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues cites this paper.

Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:38.897667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:19:57.347494Z digest=sha256:e8526ed35411ec05a43a8e2e46532e3c3fa5b61e1db11a24772c75ce1c222f29

Observation 5ee50c9b-2199-4444-828f-204af964f158 · inbound

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind cites this paper.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.252198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:09:37.399954Z digest=sha256:2262f53f9b9b3cd2736170d2252d9dcfc1c5ecde6f2b2abd0db48e4f6655a882

Observation e999a52a-b47f-41d9-9018-4f6ad92da217 · inbound

Reinforcing Human Behavior Simulation via Verbal Feedback cites this paper.

Reinforcing Human Behavior Simulation via Verbal Feedback ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:24:02.458142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T07:21:48.649289Z digest=sha256:3ede21df2efb438092ff42f946951c4c239eb2d5acf54bf3eecf2d40281fdcfa

Observation a09f3e32-ca1f-419c-bb34-e7844e7e3107 · inbound

Social World Model for Lifelong Social Intelligence cites this paper.

Social World Model for Lifelong Social Intelligence ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:37.152307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:29:11.743627Z digest=sha256:573d987ea63026bad9c718a0c6b104bf57826311e3fada157731b7a3dad834ab

Observation a5d78b53-6416-4a3c-bd8b-7996d8f73c68 · inbound

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes cites this paper.

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.699869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:22:34.367119Z digest=sha256:543adc773ec63cddfbc4f4b87ea7e9405e1981c361c43ff93dd67e48ed4ac441

Observation 5c4dded9-82ea-4717-99f6-3212b66dfe39 · inbound

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes cites this paper.

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T10:35:09.399899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:35:09.399899Z digest=sha256:7f7b48d3e2d22c6c4dcbfc7af9b89d27faed605ab1ed18ab7284f5b9257a2d84

Observation 29f06851-5dc7-4c70-875c-5fbd2e7f7f8c · inbound

Mental World Modeling cites this paper.

Mental World Modeling ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T11:07:38.423077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:07:38.423077Z digest=sha256:840ebf93d3049bfe607b88945deb35f43a201843beb9fba328bf520f48c7ff00

Observation 1d3a5ffb-3472-4559-8c2a-53f637b92169 · inbound

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs cites this paper.

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T08:29:29.096107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:29:29.096107Z digest=sha256:25d94731c191cc9100410d2a7cac0580f73d7cad0dde28048df6680a0e17504e