Pith. sign in

Paper Citation Record · LEDGER

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

As of 15 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2506.20664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20664 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:40.761038Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:40:54.002702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:46:29.064667Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb22f23c-b67e-4da7-9faa-36e978f4d73e · outbound

This paper cites Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.942223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.942223Z digest=sha256:e19d8591de103b9c29a8a36e650ed0218ca9cab6fd353fa7d09afcc4255e4668

Observation d6d5c370-f0aa-464e-8925-490fc5b23a75 · outbound

This paper cites Two” refers to “two dimensions.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Two” refers to “two dimensions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.238004Z digest=sha256:b0fbf01bd72f7a395c7f453cf0c8abf6ee9d3d66af85af078095acf662c4d728

Observation a3773a5a-8829-416f-ba64-03bd18f3fc64 · outbound

This paper cites OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OvercookedV2: Rethinking Overcooked for Zero-Shot Coordination

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.289724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.289724Z digest=sha256:eafb45db31983d5f9eeca81e6f840e1d3feee1bb954d539ec848c88ec388d12b

Observation 289c7480-7e28-4eac-a0f7-0950eb8a425f · outbound

This paper cites jazz fusion.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind jazz fusion

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T22:49:42.045426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.452292Z digest=sha256:949c64fd78992d09ae56f0b20846b6166a185da23a1b9af20c37d552829e8945

Observation d99fafda-3e55-4188-8842-b0537ab993a0 · outbound

This paper cites HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.482038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.482038Z digest=sha256:bd2490ee5950935f288db5ef6bd239382372f2aca4b89dd5fda99c3ad48ace76

Observation f391e950-a249-4a4a-b8e9-882df673023d · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Cannot Self-Correct Reasoning Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.615228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.615228Z digest=sha256:c27e3877afe66239ad32784d7e314a74afe6631ba24592905068cc8b0aa0469f

Observation f75ded91-df43-49bb-96a5-1ad767420178 · outbound

This paper cites OpenAI o1 System Card.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.679533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.679533Z digest=sha256:5ae78b0ba7d75380b15309cd14cbe2790c78cd723bdf1dd4b4d21d4cfff092fd

Observation 3235a2f6-8400-4628-8870-0aa352fd444f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.741025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.741025Z digest=sha256:47aba51abe33ebcd7e867068f081df9feac7ccb2176379b89b17a4db991cc94e

Observation b5176ce3-a507-4e6e-93cc-7aa339c8e6fd · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Evaluating Large Language Models in Theory of Mind Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.859398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.859398Z digest=sha256:9d8f0f97306aa636a9feb2d3464120b57c8a08144896529cce576d46c49e36ba

Observation 97c631b9-891d-44af-b148-69a919d4cee5 · outbound

This paper cites Revisiting the evaluation of theory of mind through question answering.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Revisiting the evaluation of theory of mind through question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.342665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:38.925060Z digest=sha256:c9c75b899719141cd67e0efc7f946732a3b8ad52fc3074e16056512599a6de5d

Observation 833347cd-9d40-43e6-a463-252e243985fd · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 17

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:49:38.986397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.986397Z digest=sha256:dfdc703aca5529c87217946cad312ece90a8548536c49a27b0d162ff9b630485

Observation 3021ab46-3775-4c1e-a7e9-fbd06809d3dd · outbound

This paper cites Avalonbench: Evaluating llms playing the game of avalon.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Avalonbench: Evaluating llms playing the game of avalon

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.331673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:39.053674Z digest=sha256:7f71727c5ddaf53e169dbd9a23eca9f915f8281eb155bed4685aa5d00b23b793

Observation a2f71c15-68d1-4123-8678-1327b178039e · outbound

This paper cites LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.116735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.116735Z digest=sha256:e6c85905b68e4d0f77107fe6806cf9f4a6caaedeab7f589ce3985240d5507615

Observation ecf6de51-8157-4b82-aacf-851b40a8b732 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Efficient Estimation of Word Representations in Vector Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.189471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.189471Z digest=sha256:ca816382ff714974b6db83ee2946a96d8e7f9bb50f5528fc6cde0bba8ffd361c

Observation 61b6a9dc-5a03-47bf-8815-3768923acaf9 · outbound

This paper cites Modeling Cross-Cultural Pragmatic Inference with Codenames Duet.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Modeling Cross-Cultural Pragmatic Inference with Codenames Duet

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:41.096852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:39.355554Z digest=sha256:3bb2e54205ae00e76937cf986ee90f71dea1bf20d42eff152313f2a145d081f0

Observation dd2ea97c-aa52-4054-999b-1fa2e8176f5e · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.387074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.387074Z digest=sha256:a54cf16f2d21e25a985d15a408091975bba401802eb0f7c72700fef87ad25391

Observation 4417e116-0091-4294-81d0-0ecffa310e8f · outbound

This paper cites OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.672207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.672207Z digest=sha256:3ab992e02f213374364704be8804b9188568199cd6c0abfccbd5d0651cc737f5

Observation 19ee2e07-e080-4e34-a782-39e510b0de30 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.847874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.847874Z digest=sha256:671767dda25fb65a506d346cad08bc57b301d5cfea293420dd7f6ec6dd050883

Observation 2a527933-1989-482c-8436-6261da33ad6e · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.921570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.921570Z digest=sha256:e5fabd480c9c5b2946be380d7dcdcd6eefae86cd788960bb29e437140b1c02d7

Observation 2522899f-71aa-4abb-bf4d-b1afc81d97d8 · outbound

This paper cites meaningful.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind meaningful

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.311032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.034128Z digest=sha256:569cfd894695d177f60734aa5468b80e1cf6e3071005f56ba0f3d9db6a9de71c

Observation bc856128-2d0c-4d08-80ff-d69e452abb36 · outbound

This paper cites Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Finally, we believe the study of pragmatic inference in LLMs to be a promising avenue for future research, which is made much easier by the release of our benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.299944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.149128Z digest=sha256:9ae21379ee8b7b0a7c0001b162b512ec198088b5e5c98d4159ee6a3c66c35c81

Observation 3f3f68de-77a2-449d-a0ce-2c972e7e425a · outbound

This paper cites Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Therefore, we set generous token limits (between 750 for non-reasoning models and up to 10000 for reasoning ones) to prevent cutting model generations prematurely

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.218772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.300874Z digest=sha256:06eac42fafea034493a791fa88215824cabb6a176089afd80bdfcbce2604c591

Observation 742e49c8-c70d-4693-bd42-9364f93b3946 · outbound

This paper cites This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u)).

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind This makes Alice’s utility U (u, m) =β log PLit(m|u) +ε log(1 − PEve(m|u))

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:41.972989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.662165Z digest=sha256:45f34a88fc60cf483abdcdef2e7e04fb0ea5561b171028cc7250d890a2931386

Observation de15b68b-1371-416d-8866-1febe25a8d03 · outbound

This paper cites airplane.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind airplane

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:41.731822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:40.761038Z digest=sha256:e0466be40128cd376b17faf41826b56b612b50bf7ab502e15b6756d3585ecf21

Observation 6256f2ba-97f7-454f-87f1-2dd5eee2ab85 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.360777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.360777Z digest=sha256:104266a2ab81219d6bc5a4d4790164b2d8214cb1582f7c9662822a35dbcca6e0

Observation a122cc52-9969-4734-9b7d-e596807e2cf7 · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.394970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.394970Z digest=sha256:9fb84e59f0f01f9cb77e5bc08691fd24cf9317260e138cdf340905034e8d9e2d

Observation 7df62d93-307a-42ec-83b8-4f1ecb67a7a0 · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.274909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.274909Z digest=sha256:1892d050b9eef4c0ff182ae264a02ccbd88d60bdc6e2be709f782fdfe7262c35

Observation 432f8f1b-756b-41a5-8a1f-b2fbae628669 · outbound

This paper cites GloVe: Global vectors for word representation.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind GloVe: Global vectors for word representation

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:42.321600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:39.230130Z digest=sha256:1b17c25d357c069c7898773cc62377649ec254fa7cabb25ddeb79c88e75ccbaf

Observation 82b6199b-6fc3-4418-85b2-4319bb337cba · outbound

This paper cites FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.805574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.805574Z digest=sha256:c9be4200f89be481cbd11f9dec2cda02e87c18b45ca57d1c4ea3cd840f1549e0

Observation 9c87e2ad-028a-42c4-bc45-db1759809ecd · outbound

This paper cites Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.391218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.391218Z digest=sha256:6fdd4f51f263ebb7e30bae6aed3f9edd6afd9c7444aba36ed1814b6127916b14

Observation 769e23d3-8007-4c0c-9f69-fb8d8f407855 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.149003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.149003Z digest=sha256:ccc2345f4fed68ec7bb725182bba8c5dc78235040ceb17c62b3d93d744ca1fa4

Observation 4e7260fd-7c34-4e43-9e8a-116036605cd2 · outbound

This paper cites ToMBench: Benchmarking Theory of Mind in Large Language Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind ToMBench: Benchmarking Theory of Mind in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:37.997496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:37.997496Z digest=sha256:a0b776dd48b869e73a96fe0882d5ca8de0a0d6be11d369b607075893d499f1df

Observation 5ddcaa60-af84-408e-804c-f494c3de74ea · outbound

This paper cites BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.475102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.475102Z digest=sha256:a76fd75747a7081e94e529b83a5aa6392c83d11518d7561bb8ba40717ea55d1c

Observation 4046dfb5-df76-4b1f-a094-fb49f729f383 · outbound

This paper cites Re-evaluating Theory of Mind evaluation in large language models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Re-evaluating Theory of Mind evaluation in large language models

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:41.405397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:49:38.555674Z digest=sha256:f04d430e468575939f48a44ece0e4f45b9bf6f61ccfca9ac39b927713e79effd

Observation e4e92b50-ee5d-41b7-9a75-9af4513427da · outbound

This paper cites The Llama 3 Herd of Models.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.231197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.231197Z digest=sha256:15e467d9aa10d185842ab917d09fcfc925ebb937fcbf57f8840cd3e977a4810b

Observation b5de8759-a0ab-40a2-a4f3-504c9e5e3919 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.076422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.076422Z digest=sha256:ece02775dad05d4a6bd0b3bd63286852069119bad901e88063acd2455be8c643

Observation 05915a83-1677-49b5-b8ec-f2b9f5f4329b · outbound

This paper cites Embodied LLM Agents Learn to Cooperate in Organized Teams.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Embodied LLM Agents Learn to Cooperate in Organized Teams

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:38.421893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:38.421893Z digest=sha256:068a37de8121af91d45a0eaf1c8c753d246772a5e996fc0fd65847628e9f1d90

Pith citing papers

Observation 04294d1a-dadb-4d38-a3a4-e58855b83b8f · inbound

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs cites this paper.

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.066605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:37:45.062718Z digest=sha256:8c4e7e50dcd87d61335d0f80240006f644908605c9b867c5873b20be0f3d49e5

Observation 8c736d73-9a8d-4b1d-897b-51f87771957c · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.730832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:2514a99b4b0d4f79a245dacd40def3b8a5bbe27ba8651a64fc8b684aa1a14615