Pith. sign in

Paper Citation Record · LEDGER

OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.06044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.06044 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:28:26.516974Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:28.198480Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e0acd5bb-3522-4a1c-a658-4ccc7d573a77 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:26.516974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:26.516974Z digest=sha256:b3d504aaa02db3a69e1c9285eeaa7edaa46e08c348ae96f07e2febabddaf130a

Observation 4417e116-0091-4294-81d0-0ecffa310e8f · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.672207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.672207Z digest=sha256:4c90a3477cfbff262056c886d78aef257844f386041512b00fc631750aa2fab5

Observation eebda44b-52d5-4a41-b0a2-000856dd5750 · inbound

Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study cites this paper.

Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:43.772624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:54:43.772624Z digest=sha256:6f99179431baf3ab1a3e4749c9c9f579f5caef7bd03f711eafc68aa1ac66e08a

Observation 55a8ea9b-0b88-4484-843e-cf715432249c · inbound

H2HTalk: Evaluating Large Language Models as Emotional Companion cites this paper.

H2HTalk: Evaluating Large Language Models as Emotional Companion OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:29.418863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:29.418863Z digest=sha256:4fc046c55c8c65506ce508d9d71fe60c71130e6b1a303080ca74ffbfb4c23bc5

Observation 32fb6eeb-3b72-4e89-99b5-17558e1fce13 · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.320417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:7b49812e19ea9cce1bbb632b50ca2b3259d5bef090e9010d1cd1d60d69358e38

Observation 2da6e9a5-8134-4b7d-ba6e-194b5081b3ac · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:28.199822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:7bb3db26be4f38968f9124f6ca658cb1a783d03ce1197d305bcbbe312fd6ab1a

Observation 4604e8a3-029a-4c0c-9f0a-3b24157efb5e · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.654373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:b6a8d1224d06deabd941c83a56ecc68eba6ddd2cd3fbaea380b8bd72c7b5e93e