Pith. sign in

Paper Citation Record · LEDGER

OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.06044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.06044 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:28:26.516974Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:28.198480Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e0acd5bb-3522-4a1c-a658-4ccc7d573a77 · inbound

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models cites this paper.

From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:28:26.516974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:28:26.516974Z digest=sha256:b3d504aaa02db3a69e1c9285eeaa7edaa46e08c348ae96f07e2febabddaf130a

Observation 4417e116-0091-4294-81d0-0ecffa310e8f · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.672207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.672207Z digest=sha256:4c90a3477cfbff262056c886d78aef257844f386041512b00fc631750aa2fab5

Observation eebda44b-52d5-4a41-b0a2-000856dd5750 · inbound

Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study cites this paper.

Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:43.772624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:54:43.772624Z digest=sha256:6f99179431baf3ab1a3e4749c9c9f579f5caef7bd03f711eafc68aa1ac66e08a

Observation 55a8ea9b-0b88-4484-843e-cf715432249c · inbound

H2HTalk: Evaluating Large Language Models as Emotional Companion cites this paper.

H2HTalk: Evaluating Large Language Models as Emotional Companion OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:29.418863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:29.418863Z digest=sha256:4fc046c55c8c65506ce508d9d71fe60c71130e6b1a303080ca74ffbfb4c23bc5

Observation 32fb6eeb-3b72-4e89-99b5-17558e1fce13 · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.320417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:597a8d90df981f726f474b228659e8b3b78dac2c1f2de3296a37e466416fb2f5

Observation 2da6e9a5-8134-4b7d-ba6e-194b5081b3ac · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:28.199822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:099e7e518b185a98efbe7942aed08d608379f688bf0c517910c77024fe9b82bf

Observation 4604e8a3-029a-4c0c-9f0a-3b24157efb5e · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.654373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:9df8d92e17e36b8db62f163380cdb87451f64987054f0c92b784a456ebc7115b