Pith. sign in

Paper Citation Record · LEDGER

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2506.20664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20664 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:40.761038Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:40:54.002702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:46:29.064667Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb22f23c-b67e-4da7-9faa-36e978f4d73e · outbound

This paper cites Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.942223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.942223Z digest=sha256:8dc6479f3249b3c83e7f0f62c394995fccb774857105a579549a4895be6e9e8d

Observation d6d5c370-f0aa-464e-8925-490fc5b23a75 · outbound

This paper cites Two” refers to “two dimensions.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Two” refers to “two dimensions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.238004Z digest=sha256:b02e3785abf03a8999040f0120f51d9f15952426ccce45f2b34737182c35c519

Observation a3773a5a-8829-416f-ba64-03bd18f3fc64 · outbound

This paper cites OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.289724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.289724Z digest=sha256:2d1df0161c21e369672c95f5ad054acbc48b4d04fd2433e1d59a91ab8c42f232

Observation 289c7480-7e28-4eac-a0f7-0950eb8a425f · outbound

This paper cites jazz fusion.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind jazz fusion

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T22:49:42.045426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.452292Z digest=sha256:c66a7102187a6487a84c8b808b343440c2a941c9fb02977cc05390600200573b

Observation d99fafda-3e55-4188-8842-b0537ab993a0 · outbound

This paper cites HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.482038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.482038Z digest=sha256:e7c81416ec15d60a7d2735706c910999b32c97278132bf4d1fcc57bd85d2edbd

Observation f391e950-a249-4a4a-b8e9-882df673023d · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Cannot Self-Correct Reasoning Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.615228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.615228Z digest=sha256:bccdf3b1f6ba3f6a908c6fbbe6a1661658de8b8d342d5fe4a0d04927f2c961d6

Observation f75ded91-df43-49bb-96a5-1ad767420178 · outbound

This paper cites OpenAI o1 System Card.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.679533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.679533Z digest=sha256:9aa3d9e0f8ae5b856d6cfc39269ea1035b255c355d4dd7e3b9001f0c23439719

Observation 3235a2f6-8400-4628-8870-0aa352fd444f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.741025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.741025Z digest=sha256:1ad2353390475e3654df2707e0241cb56dd9affa32d0c9ec33b99f103815421b

Observation b5176ce3-a507-4e6e-93cc-7aa339c8e6fd · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.859398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.859398Z digest=sha256:a96d14d86a23ae3b22037b2666d0e3b05c4191d910e750139c73e10db3934f8f

Observation 97c631b9-891d-44af-b148-69a919d4cee5 · outbound

This paper cites Revisiting the evaluation of theory of mind through question answering.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Revisiting the evaluation of theory of mind through question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.342665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:38.925060Z digest=sha256:d02ae9e69a7171ab62194a43aafe3539ae092b36b7117a2b65222abb9d6d05ed

Observation 833347cd-9d40-43e6-a463-252e243985fd · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 17

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:49:38.986397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.986397Z digest=sha256:375d30806d5663eb9246db7bf7f0501870c6b51ac04f84762a71e46ab6e4491a

Observation 3021ab46-3775-4c1e-a7e9-fbd06809d3dd · outbound

This paper cites Avalonbench: Evaluating llms playing the game of avalon.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Avalonbench: Evaluating llms playing the game of avalon

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.331673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:39.053674Z digest=sha256:e0551a7d88c1b19b1387e7af10a8be33931760f4623494f805b81a853075b97d

Observation a2f71c15-68d1-4123-8678-1327b178039e · outbound

This paper cites LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.116735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.116735Z digest=sha256:50cfb5cb2f9fc3da800251ed93a97c1beb9c5730101eb73fa776958733be764b

Observation ecf6de51-8157-4b82-aacf-851b40a8b732 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Efficient Estimation of Word Representations in Vector Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.189471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.189471Z digest=sha256:8f7c35670033c5897e90ddabfc5eb9524ad01fa827c3995bf3ca6b1dfe237eab

Observation 61b6a9dc-5a03-47bf-8815-3768923acaf9 · outbound

This paper cites Modeling Cross-Cultural Pragmatic Inference with Codenames Duet.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Modeling Cross-Cultural Pragmatic Inference with Codenames Duet

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:41.096852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:39.355554Z digest=sha256:1f2afa16b0b9839a4cd6ecd66431cf6cb2ce19826c84202f8ee2a230a0f27402

Observation dd2ea97c-aa52-4054-999b-1fa2e8176f5e · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.387074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.387074Z digest=sha256:9e32b47464fccafa6db1f9d4dfdfaf368fd2ed5034026bcedab1e3cd37e64375

Observation 4417e116-0091-4294-81d0-0ecffa310e8f · outbound

This paper cites OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.672207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.672207Z digest=sha256:0c6634df30702bfb9105ad945f775c4644e3d4f9e104a5659cc26874ed56ad2c

Observation 19ee2e07-e080-4e34-a782-39e510b0de30 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.847874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.847874Z digest=sha256:51c73dd7af414b0a4dfb6f91c303983b66bf641c8e5e213ae3de6894cdc51cc1

Observation 2a527933-1989-482c-8436-6261da33ad6e · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.921570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.921570Z digest=sha256:e5d09735ce7581b2409edb39c67a1dd6193352bb29343f9e9d05336b5e2422c3

Observation 2522899f-71aa-4abb-bf4d-b1afc81d97d8 · outbound

This paper cites meaningful.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind meaningful

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.311032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.034128Z digest=sha256:3eeb39161363308065e950ea537721711748fa36ccd34d9998ae2a7ba4adfb29

Observation bc856128-2d0c-4d08-80ff-d69e452abb36 · outbound

This paper cites Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.299944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.149128Z digest=sha256:c7a467dff5a848187163c4ad2b086c7d3ce203643ebab64621eb839fbd03a05b

Observation 3f3f68de-77a2-449d-a0ce-2c972e7e425a · outbound

This paper cites Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.218772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.300874Z digest=sha256:75f4a2b663b6e04b1a974edcce3d8900f95858e9d91a64b90ac9abc386e1577c

Observation 742e49c8-c70d-4693-bd42-9364f93b3946 · outbound

This paper cites This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u)).

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u))

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:41.972989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.662165Z digest=sha256:c221ee9291db0432ca20ece1bcea608179e4e591da3ecccf11a5b2db6c3c9ab6

Observation de15b68b-1371-416d-8866-1febe25a8d03 · outbound

This paper cites airplane.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind airplane

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:41.731822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:40.761038Z digest=sha256:ed125b2599b4a9c91c7a80bfa58f2bc5dba94424f9bcf534afc91472db076e38

Observation 6256f2ba-97f7-454f-87f1-2dd5eee2ab85 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.360777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.360777Z digest=sha256:26fa5b1f96aafd740e1cf2b67a8d625710554a4fbfbad9ffaf4bc04750b01553

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:721860dbf493a7a2d5f64abe647c29b053cc5ef5fa8b172de2671340ce7c0d30

Observation 7df62d93-307a-42ec-83b8-4f1ecb67a7a0 · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.274909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.274909Z digest=sha256:d7c197313f6070e943d0bacda8e8ea5ffe9643ef7afff773f704b7b6fb3992b1

Observation 432f8f1b-756b-41a5-8a1f-b2fbae628669 · outbound

This paper cites GloVe: Global vectors for word representation.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind GloVe: Global vectors for word representation

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.321600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:39.230130Z digest=sha256:63eed5cfdabb799212246d28d8116bbf88a6f20e0df54fd2d08e337694153767

Observation 82b6199b-6fc3-4418-85b2-4319bb337cba · outbound

This paper cites FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.805574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.805574Z digest=sha256:5bd87ec9fcc738eefeeee4e4469a15a6c15604e1b9d463ab9d77e6bcab35a4f1

Observation 9c87e2ad-028a-42c4-bc45-db1759809ecd · outbound

This paper cites Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.391218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.391218Z digest=sha256:e106deef41891244b35946565e002d4e035704bc88255a8d7af00c164f813d75

Observation 769e23d3-8007-4c0c-9f69-fb8d8f407855 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.149003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.149003Z digest=sha256:056f08e2486b140873f0e235d2757c59af01b3f86a0812cbdad78e62f4feacc4

Observation 4e7260fd-7c34-4e43-9e8a-116036605cd2 · outbound

This paper cites ToMBench: Benchmarking Theory of Mind in Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.997496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.997496Z digest=sha256:543f40934fb6121116486d732f2f865d3828fd1533a5a6341688dd591db08fe6

Observation 5ddcaa60-af84-408e-804c-f494c3de74ea · outbound

This paper cites BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.475102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.475102Z digest=sha256:8a2b926ee46dbdd7a81b6fdda38409d468b49ec7f352d005f093b3a5a549438f

Observation 4046dfb5-df76-4b1f-a094-fb49f729f383 · outbound

This paper cites Re-evaluating Theory of Mind evaluation in large language models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Re-evaluating Theory of Mind evaluation in large language models

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:41.405397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:38.555674Z digest=sha256:c08b30a8350084b2ebbef7296a086330163c5884b984e36eb13459a2a8679066

Observation e4e92b50-ee5d-41b7-9a75-9af4513427da · outbound

This paper cites The Llama 3 Herd of Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.231197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.231197Z digest=sha256:8647f24c5c53db2f255536eee67c58820b0f567d8d44c329e4aabc3886851a41

Observation b5de8759-a0ab-40a2-a4f3-504c9e5e3919 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.076422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.076422Z digest=sha256:46a35227a08d3d71f5debf963d6b9ed7e5cb2bc380844948e3a54d5d8dbc36f1

Observation 05915a83-1677-49b5-b8ec-f2b9f5f4329b · outbound

This paper cites Embodied LLM Agents Learn to Cooperate in Organized Teams.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Embodied LLM Agents Learn to Cooperate in Organized Teams

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.421893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.421893Z digest=sha256:8f3ededf8155d2d23816a31ac9ced256eda9ba7824864594060f90905a262f9a

Pith citing papers

Observation 04294d1a-dadb-4d38-a3a4-e58855b83b8f · inbound

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs cites this paper.

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.066605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:37:45.062718Z digest=sha256:e955db1d0221e08142f559cf8c3a1732c123fa6a9d1da3d35fad259fe4733adf

Observation 8c736d73-9a8d-4b1d-897b-51f87771957c · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.730832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:b80d589b9dc213eaea26fc92b5648b6b441709a2b6b0e28459765dca44d8a6cf