Pith. sign in

Paper Citation Record · LEDGER

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 100 of 164 outbound references and 7 inbound Pith citation observations for arXiv:2505.23713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23713 v1

Coverage vector

measured 100 of 164 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:21.376172Z

measured 107 of 107 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:16:48.882711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:44.817622Z

Reference resolution

100 of 164 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c19d23c-e099-4afe-8365-8597d00d0577 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.Advances in Neural Information Processing Systems, 36:59662–59688, 2023.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks.Advances in Neural Information Processing Systems, 36:59662–59688, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.560836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.560836Z digest=sha256:c81073e0f22e1e173352ed21cada3b51433a3bd09dcceee1ba2b8de2ae0f590e

Observation c07aa0af-befa-41b6-a881-62f4cba94803 · outbound

This paper cites Unveiling the power of language models in chemical research question answering.Communications Chemistry, 8(1):4, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Unveiling the power of language models in chemical research question answering.Communications Chemistry, 8(1):4, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.641139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.641139Z digest=sha256:0f4f8426998fb7778aefdef56832aa7354f6d4531d13aac6d002e899e968e5b3

Observation c93a395a-dc3f-4c91-8ab2-24335af568d2 · outbound

This paper cites Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.681909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.681909Z digest=sha256:befe67b505dae89b05091b36d206abd5ffd33f991400f80a6b5bddf288849c07

Observation 10d83bd2-6924-4e6f-a2db-8a76e44650ed · outbound

This paper cites Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.Nature Communications, 15(1):5649, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.Nature Communications, 15(1):5649, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.787564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.787564Z digest=sha256:11809495fc2ebd108c066a5b4ebfe7ddb9ad5386d5e619be64e800010cdbc0d0

Observation ef91d5c4-f0e5-44ea-96bf-9ae94dfc3a25 · outbound

This paper cites Evalu- ating and mitigating bias in ai-based medical text generation.Nature Computational Science, pages 1–9, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evalu- ating and mitigating bias in ai-based medical text generation.Nature Computational Science, pages 1–9, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.877179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.877179Z digest=sha256:e04dec23fbbf307299132f1feda0ad6cd4f886b017f8f7baf3c004a58f7baa88

Observation 73aa2aa5-0385-4117-b5ef-d1edfe8096e6 · outbound

This paper cites Llm-mod: Can large language models assist content moderation? InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–8, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llm-mod: Can large language models assist content moderation? InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–8, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.952094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.952094Z digest=sha256:c2134ee73cf0d922ed4a16583907aa351b684531100fd2e6269b3b021aad838d

Observation 0214642d-ba8c-4bd2-8a1e-d138d1e93ce2 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.036015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.036015Z digest=sha256:d73996de2f195479d335092d8f71275fe6fd3a8f6c8feea61ebe644abab911ff

Observation 87e8edcc-6e5b-41f2-8c4f-92e4b614862b · outbound

This paper cites Scaling up llm reviews for google ads content moderation.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Scaling up llm reviews for google ads content moderation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.119173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.119173Z digest=sha256:6f090c61d0918a4d007717890b38ffb613c18f77692ef4f45cb0d766ce7d8c72

Observation c0b1784b-7f58-45fa-b2fa-a2a1da2ac8bd · outbound

This paper cites Hate Personified: Investigating the role of LLMs in content moderation.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Hate Personified: Investigating the role of LLMs in content moderation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.169211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.169211Z digest=sha256:8a14b5f0624736b444130bd84277850efd3352e9b97830650802c98d86dd3603

Observation 6ae55d95-54c4-4bc4-9ce5-a2aefbd9603e · outbound

This paper cites Autonomous agents for collaborative task under information asymmetry.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Autonomous agents for collaborative task under information asymmetry

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.239780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.239780Z digest=sha256:50533fcce9513e32c17c814576975f636313303212fd9576018901d3ac0360a7

Observation e7a85e38-eebe-4b4f-b1f6-b7b7a774a32a · outbound

This paper cites Halc: object hallucination reduction via adaptive focal-contrast decoding.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Halc: object hallucination reduction via adaptive focal-contrast decoding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.374316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.374316Z digest=sha256:bfa0506a1db4ff174f34a674be0de1fdc8cc1880c35c361d4ce37e9969505ff0

Observation c8ecb7a4-1bdb-4242-8599-1271f87ff4f9 · outbound

This paper cites LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.486260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.486260Z digest=sha256:48b7d317e69dcd0701a018414a36569cf1d8481c49f3d6ea83fde20d4b5f09c5

Observation 5b54ee77-7f94-4d69-9951-7ae33443b3b5 · outbound

This paper cites The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.630040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.630040Z digest=sha256:fe111ff8b85951f57df4f61be1433898970cd9743d4d3798d01a30ec634440ba

Observation 64aad1e1-38c5-4a3f-b89d-1146e8187355 · outbound

This paper cites Decoding echo chambers: LLM- powered simulations revealing polarization in social networks.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Decoding echo chambers: LLM- powered simulations revealing polarization in social networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.698038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.698038Z digest=sha256:5a849496b9724c46d7a46c8750011277b22a43293a15da2dfdcd9f2d1dc41f86

Observation 8eebf24c-cb28-4630-a5e7-9ac7f70642e6 · outbound

This paper cites Safewatch: An efficient safety-policy following video guardrail model with transparent explanations.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Safewatch: An efficient safety-policy following video guardrail model with transparent explanations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.850117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.850117Z digest=sha256:1b326897119e971e6a0f216a3a7bfec2624ee31c9ed8b7b665ae89624e689a54

Observation 53d70317-cbd7-4854-a6df-03a57914247f · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:11.951679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:11.951679Z digest=sha256:678404fa40c5ef67ede08a93f0b28a29140584d3da5b9e0248cfd1d4bb16cad3

Observation 9d09d512-727a-42b6-9f28-256481b93c14 · outbound

This paper cites Exploring Large Language Models for Word Games:Who is the Spy?.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Exploring Large Language Models for Word Games:Who is the Spy?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.071230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.071230Z digest=sha256:e90024ec52f8c0f76c7cbb3439fc8e73ecd688652a344a9dbd4ecbc393b17dd2

Observation 2cfbf2c1-199c-4e48-a2b9-febad1800e59 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Chatbot arena: An open platform for evaluating llms by human preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.176186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.176186Z digest=sha256:4d1b251eaff2c120907e7642e82d4401bda2b7d4294bbbdb5561285d7c8c3b5d

Observation 5bba9b19-3104-4b81-8280-0cfa30e6eb40 · outbound

This paper cites On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.276453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.276453Z digest=sha256:7cdb4025bd0da226b579a4ecab2f259ffd90bece5b26f8d038532d76acd58251

Observation d9f38520-a65f-4128-a3f0-10133692af8c · outbound

This paper cites Trusteval: A dynamic evaluation toolkit on trustworthiness of generative foundation models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Trusteval: A dynamic evaluation toolkit on trustworthiness of generative foundation models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.374411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.374411Z digest=sha256:2c705b8388bf955937a50b65a0e7307f7664c9c40ecde74514ed856e429310fc

Observation c47236ff-3e92-407a-b1a9-6bdb42ef4fcc · outbound

This paper cites Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evaluating large language models in theory of mind tasks.Proceedings of the National Academy of Sciences, 121(45):e2405460121, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.480039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.480039Z digest=sha256:f59ac80426f34c663de8e5a5006ecd2141b4ff239ef3cd33feb539ffdf897da1

Observation 1861385d-1771-442f-b1f7-b6ffd1243484 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.627945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.627945Z digest=sha256:fa131c4a03a26225eede182c5e6d5f5f243d6fa68d42a9fb9c7e34a769169fa4

Observation ced0f817-bd1d-4b6e-8ce2-9bc40924d527 · outbound

This paper cites Cross-lingual pitfalls: Automatic probing cross-lingual weakness of multilingual large language models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cross-lingual pitfalls: Automatic probing cross-lingual weakness of multilingual large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.786418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.786418Z digest=sha256:308499629fee681ca02f553dbd5e9ce64226db3d6835a8c423a9908fdef63739

Observation 144b460c-365a-4c3b-9346-af7033c27931 · outbound

This paper cites Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:12.924587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:12.924587Z digest=sha256:e41df1f0ac05f2c14f9051b46c2e557a20983fcf53e063098dfeb683c5091418

Observation a0ce336a-c6c3-40bb-a732-555597637bcf · outbound

This paper cites Honestllm: Toward an honest and helpful large language model.Advances in Neural Information Processing Systems, 37:7213–7255, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Honestllm: Toward an honest and helpful large language model.Advances in Neural Information Processing Systems, 37:7213–7255, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.044373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.044373Z digest=sha256:4eddd7f7f195fdb05d69ce04b78998a0606a0f3c46f7ff988d858d8b515196a7

Observation b705d457-b9ad-4bf6-aedb-6512f0c9a1e6 · outbound

This paper cites Vuldetect- bench: Evaluating the deep capability of vulnerability detection with large language models, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Vuldetect- bench: Evaluating the deep capability of vulnerability detection with large language models, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.160590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.160590Z digest=sha256:9f4911c1736c53446df1e2f682e85a58806af322a2df36aff900fcb59b19ef2d

Observation 53675f87-0002-40fc-ac33-6f073a12c82b · outbound

This paper cites Datagen: Unified synthetic dataset generation via large language models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Datagen: Unified synthetic dataset generation via large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.306776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.306776Z digest=sha256:4f339b809cb09fb70e82a22228e22c8d7c212f10e421bc095765bd1042903966

Observation cb390568-9bb4-43d4-9110-5414a5a2fd84 · outbound

This paper cites Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.459963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.459963Z digest=sha256:94583994026cadb732ffb5568081c81c9876914f889fb49c558c807ed580bfce

Observation 6c5fecf5-1d5a-4114-beb0-125e884bc2b3 · outbound

This paper cites Nqe: N-ary query embedding for complex query answering over hyper-relational knowledge graphs.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Nqe: N-ary query embedding for complex query answering over hyper-relational knowledge graphs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.575817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.575817Z digest=sha256:a97c24db66705e7291448d17d80f9986664802825fa0c42c7aaa405510afc3ea

Observation 4c05919c-aa1e-4ef6-8b37-46ce0ca263c2 · outbound

This paper cites Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.720220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.720220Z digest=sha256:7dc6170e4e9a648ef6062d316eace75035203bc95842b347c10f32d2c01362e5

Observation 68b15ff9-d70c-4251-b8c0-e09a4dad781e · outbound

This paper cites Gta: Graph theory agent and benchmark for algorithmic graph reasoning with llms, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gta: Graph theory agent and benchmark for algorithmic graph reasoning with llms, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.834447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.834447Z digest=sha256:2f7e52946dfbb9e4b5fbe7c82e78ded42221c6ede76de2186ef6a6e2d951c981

Observation 96fdb768-209a-4ab6-be35-47b046e434cc · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models SocialIQA: Commonsense Reasoning about Social Interactions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.983646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.983646Z digest=sha256:469f1e9a32c01ef87a33e2bdcc5306d35cf88351376064084b9f32fb0f76d7d8

Observation 2caac88c-83d4-43a6-b5e8-d34fda8f6c67 · outbound

This paper cites Atomic: Anatlasofmachinecommonsense for if-then reasoning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Atomic: Anatlasofmachinecommonsense for if-then reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.142371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.142371Z digest=sha256:e3cd77a58493e82aced0cfd2069a1dd299ef39b39a0563e06d2f4bcdbe3f7b5e

Observation 562d762f-b506-4b2a-93db-7968ba8248e3 · outbound

This paper cites CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.262662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.262662Z digest=sha256:ab09f61d7268e7999c3163cdba6cd6d085c8a6d615fa16bb24b7f748e58f3e43

Observation d4eb93cc-e80b-4bc4-b8cb-90455851fc98 · outbound

This paper cites GoEmotions: A Dataset of Fine-Grained Emotions.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models GoEmotions: A Dataset of Fine-Grained Emotions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.387941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.387941Z digest=sha256:7e5b8d6c7b01f2851bd10854a4666fc3f8e38d5d1297b64ad3b65bbf494d7834

Observation 892c49eb-0b48-4ca7-ba06-f328200e6c82 · outbound

This paper cites CommonGen: A constrained text generation challenge for generative commonsense reasoning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CommonGen: A constrained text generation challenge for generative commonsense reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.537634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.537634Z digest=sha256:1df07fe0876a05d1ca4919fa1e80c78aed8437815eb723a1f0e8ae09ea9d87d1

Observation fae1bdab-2a69-466e-af42-439d37bc5b3a · outbound

This paper cites Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.673881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.673881Z digest=sha256:0a83da82996154a57e99fdbede1cf8a4f21837bcb6d3607046ee480785b017bc

Observation 3f772664-b5eb-427f-8021-856e067cfbc8 · outbound

This paper cites Evaluating Theory of Mind in Question Answering.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Evaluating Theory of Mind in Question Answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:14.837395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:14.837395Z digest=sha256:8b8db7132636e2cb58f84f18af77a92392424b4a7d5ec8c5f51b3c9148521ef7

Observation d50980b8-64d9-48ee-b6e8-489b317aa38a · outbound

This paper cites Social Chemistry 101: Learning to Reason about Social and Moral Norms.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social Chemistry 101: Learning to Reason about Social and Moral Norms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.026411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.026411Z digest=sha256:41208f2c187395bfe0dad999006b3939e0e713d0dec66f31d49a1dcadb832ef6

Observation b9e34c8e-6b11-4a98-bdea-0ba7aeae428c · outbound

This paper cites Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.176837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.176837Z digest=sha256:d34d87ca0679765c82307ef55ff8acff3d59481d221395f8a5fbd749b6eef0f5

Observation 5096160d-dc35-4a12-bd7d-d52ca33d0792 · outbound

This paper cites DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.847772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:15.313592Z digest=sha256:50f7e4da3b7a42615b7102f9d3e3c7167d4b2b2fc12965b3753db31796033813

Observation ed687835-2afc-4ef2-b5bd-8032534ebd3b · outbound

This paper cites Understanding social reasoning in language models with language models.Advances in Neural Information Processing Systems, 36:13518–13529, 2023.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Understanding social reasoning in language models with language models.Advances in Neural Information Processing Systems, 36:13518–13529, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.443725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.443725Z digest=sha256:c4aa50b7d906c62777fd3b3456e041ff60c462286686bea866fb6f1408f3aeab

Observation bed73b27-ddcb-4714-b0e6-b02bd592d5fa · outbound

This paper cites Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.575617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.575617Z digest=sha256:c0d4e0f9070a543a32628de8bea60747b3d3ab568f50fe029553fd8bf043589e

Observation 160bc2c0-759b-4f4d-9147-24b6198c4e06 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.723111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.723111Z digest=sha256:9333cdf256fb244d5dc2b396a7779eeeb68bdf8d7d865d04a4dc6711853ebb24

Observation 6df2bdba-089d-4aee-bafa-a3f3e4da6171 · outbound

This paper cites Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.845155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.845155Z digest=sha256:1d4a9ba8ed6f28c7b8aa940e46eb9062247b88d54a69a375b61da86400226366

Observation 8346a62d-881e-42ce-8837-f77d96954f15 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.951141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.951141Z digest=sha256:2e1cbc070b46bf69cc51d65db6225c58698802bac4273868facbd8aa314fdbb1

Observation 7afb8966-e0ae-4455-b482-01614fbc96dd · outbound

This paper cites Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.080047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.080047Z digest=sha256:5b9967d6a7eb81685a22436fa2295117511fb96bd64198f4036d3b452da72f42

Observation 2771bbea-358a-45a0-8a3c-0386c43f734a · outbound

This paper cites Oxford University Press, 2014.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Oxford University Press, 2014

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.198939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.198939Z digest=sha256:bfafe4e346093a58692d72bd31dab8d7257c9b474f093b35a8e43e6eed943f14

Observation caa6c6f4-d336-4015-bb62-7ccfabd777d7 · outbound

This paper cites MIT press, 1999.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models MIT press, 1999

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.279846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.279846Z digest=sha256:691a8a25e973cad457652f69b004b4228851b7ccb036a4b10624580df24cabae

Observation 4e9c919f-6e1e-4fc7-ba27-efe5618d6d4e · outbound

This paper cites Social cognition in humans.Current biology, 17(16):R724–R732, 2007.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social cognition in humans.Current biology, 17(16):R724–R732, 2007

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.346129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.346129Z digest=sha256:6806adc9148aff72fb4f772ab64b8dede6e109f6c6be9bbdca99dd0ba690304c

Observation e27229f5-1234-43d0-bed7-c944e88d070f · outbound

This paper cites Counterfactual thinking.Psychological bulletin, 121(1):133, 1997.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Counterfactual thinking.Psychological bulletin, 121(1):133, 1997

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.404998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.404998Z digest=sha256:fc0b394ed13600bbdc4f807c824ab183632103fc4a1f0ec810bc60b5bfefe0eb

Observation 42f63aa8-154a-4dfd-a993-13a2ecee78e4 · outbound

This paper cites CLOMO: Counterfactual Logical Modification with Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models CLOMO: Counterfactual Logical Modification with Large Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.628490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:16.459659Z digest=sha256:e4eec5fceba460e1bccbe9f7f70a3d1940b9c64c8d93c148aa23a67ed7f48152

Observation 06583048-38d9-4073-8ab5-b8fd7f349246 · outbound

This paper cites LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.514813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.514813Z digest=sha256:a28b9e6e2daa1734cc51a229519a1606edabfadfd69d732659ebfe28dd3dd6b9

Observation 57e784be-1ab8-44fe-a2cb-722ecdff9b39 · outbound

This paper cites What is agency?American journal of sociology, 103(4):962– 1023, 1998.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models What is agency?American journal of sociology, 103(4):962– 1023, 1998

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.598584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.598584Z digest=sha256:2f53834904236f66de2600c9731e7f5d1e455f0970f48d2511ed56dffa5b53a3

Observation a5ac645c-90ae-47fd-a850-83a416fdf9b5 · outbound

This paper cites Penguin, 2014.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Penguin, 2014

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.670659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.670659Z digest=sha256:c7de180f1a65e9ea706d65a6596e65070d96ceafbc6227bd64dad17fcf8b5b59

Observation 531f0aae-4831-4879-8414-49c086fa0760 · outbound

This paper cites Cambridge University Press, 2008.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cambridge University Press, 2008

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.710690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.710690Z digest=sha256:9cb051f53fb4570659113b62969df63424732723de8d33d1f56922bfa16c39f0

Observation a7d3c459-8d16-4d8c-9bac-7aa76bc442e6 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Generative agents: Interactive simulacra of human behavior

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.890152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.890152Z digest=sha256:456773c08a55bbd68ad5863e129a054ffcdb93225bad898b159ec3305d505621

Observation 34db5b5e-464e-4998-a1af-daab964ecc81 · outbound

This paper cites Shieldagent: Shielding agents via verifiable safety policy reasoning.arXiv preprint arXiv:2503.22738, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Shieldagent: Shielding agents via verifiable safety policy reasoning.arXiv preprint arXiv:2503.22738, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:16.957072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:16.957072Z digest=sha256:3084bd38f38379a500abe1b2cf338cef82e91bec517799b93b60bad1b7b4c9cb

Observation a3664a79-20f0-44d4-b9a6-75cf166c320a · outbound

This paper cites Fusing heterogeneous data: A case for remote sensing and social media.IEEE Transactions on Geoscience and Remote Sensing, 56(12):6956–6968, 2018.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Fusing heterogeneous data: A case for remote sensing and social media.IEEE Transactions on Geoscience and Remote Sensing, 56(12):6956–6968, 2018

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.044966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.044966Z digest=sha256:a5299e39898775debfdeb6f95d2211ca23d880ef65e98a1df7eabfa396542089

Observation 5da8a377-711a-4360-b89c-e46dd09446d9 · outbound

This paper cites How noisy social media text, how diffrnt social media sources? InProceedings of the sixth international joint conference on natural language processing, pages 356–364, 2013.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models How noisy social media text, how diffrnt social media sources? InProceedings of the sixth international joint conference on natural language processing, pages 356–364, 2013

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.136964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.136964Z digest=sha256:9fd192518c5d4b8ab6346409d49a68719bb1e6db25dbcff4755feb30351ab5aa

Observation c6d3cc52-1d5d-4440-a460-5bdb623a1096 · outbound

This paper cites Social-ecological systems as complex adaptive systems.Ecology and Society, 23(4), 2018.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Social-ecological systems as complex adaptive systems.Ecology and Society, 23(4), 2018

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.228510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.228510Z digest=sha256:3f5a0cd643f8aed96a38236134388aecfc8ff30a5b1bf5f6008ba6956a68480f

Observation 70afb51a-c228-4a3c-900a-458fff4c341f · outbound

This paper cites The science of fake news.Science, 359(6380):1094–1096, 2018.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The science of fake news.Science, 359(6380):1094–1096, 2018

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.336305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.336305Z digest=sha256:29b8f8288dc94ab5f66936918fc321b9a1d9d076e7741828e5144f93dacb7315

Observation f9039d21-c660-4f93-bfb5-837990d164fc · outbound

This paper cites Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases.Advances in Neural Information Processing Systems, 37:130185–130213, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases.Advances in Neural Information Processing Systems, 37:130185–130213, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.421245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.421245Z digest=sha256:38d590eaaab98cd5f214123451da54b2f626b0502d19bcaaed24f93b90b619e9

Observation 7a1f52dc-98bd-4507-af31-9d066e995cb7 · outbound

This paper cites Automated Design of Agentic Systems.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Automated Design of Agentic Systems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.499384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.499384Z digest=sha256:ce144a23059a70f2c67bbf9ce3bb7d8e38c420abded46ac54331a14f5e8c3bf3

Observation 5c3c57b8-a7c9-4502-9dbb-dea711fc0f59 · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AFlow: Automating Agentic Workflow Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.603078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.603078Z digest=sha256:e460052bc30863cfebc30d063802b5f199e30e0103ea92a3c48636f91fb934af

Observation 26de77b4-bcac-45a5-b411-d2f40aa55df2 · outbound

This paper cites Multi-agent Architecture Search via Agentic Supernet.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Multi-agent Architecture Search via Agentic Supernet

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.691822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.691822Z digest=sha256:7eb74c9c47157046c93b33a12e285055484adfc27e858a5989e02d89253d04a6

Observation c10eb10e-014a-45aa-bf32-90065f8b8b7c · outbound

This paper cites Blood on the clocktower, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Blood on the clocktower, 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.795238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.795238Z digest=sha256:5dcbc896cd6951d760694714bfb89d93a3aa002cf716e6efea4db77bdaa3f3e2

Observation b4c509bb-e103-454c-b7a1-5aa47847ca92 · outbound

This paper cites Improving fac- tuality and reasoning in language models through multiagent debate.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Improving fac- tuality and reasoning in language models through multiagent debate

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.883268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.883268Z digest=sha256:5d19f89084cd8722a84b8f8e52fefb3a8a37ac7830906c611f45ada7f12dd43a

Observation a3087291-5e04-4242-b044-f3d893e28ac5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:17.974595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:17.974595Z digest=sha256:25001787ae18f9c0fd6ec2eb5e18855355cd2063ea315314f5fa4db392d72e0f

Observation c90b2b0d-60c0-4ace-864f-e2afc45173c2 · outbound

This paper cites Dyflow: Dynamic workflow framework for agentic reasoning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Dyflow: Dynamic workflow framework for agentic reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.050316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.050316Z digest=sha256:ff7a197ef1de88eb9eb89537b2dcc1e66322e020d95b990b557737ad50035b94

Observation fdf1ecc5-6a52-428f-b5b2-d3d115afe017 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.133257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.133257Z digest=sha256:2f5e6ac63455d9c952ff9ac00d8b94bf8e674e526faff452722b21d1da3d5d57

Observation 4f5ebcac-d920-474e-bc80-35af9b5760a6 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.247145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.247145Z digest=sha256:b47f562dad6381c89f425217d211526c67cbed358f0f92d21807c99ce07891c1

Observation 3cf05bba-a277-4da6-91e2-2a1a66b57aa0 · outbound

This paper cites Glucose: Generalized and contextualized story explanations.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Glucose: Generalized and contextualized story explanations

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.315743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.315743Z digest=sha256:81acaa8c0fbea41084eca1f8f8db5e2ddc5b72ef208e4fdc4aa1869f70bcf267

Observation 9d30a6f0-1c1a-41be-95d2-18e1f0d21c23 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.418665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.418665Z digest=sha256:17c199ace9286f7f38b13e0d434ffe7fad288ef549a676181879c5f6f7657676

Observation a74cdb97-f283-4414-8db2-991dbf3c0f13 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.491278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.491278Z digest=sha256:8dc703716883bec34553ffc115d7371554eb4c689eec5f411d52fc9f063bde4d

Observation 75b2f6fc-b31c-4bdc-8509-7bafd5e238d5 · outbound

This paper cites Theory of mind may have spontaneously emerged in large language models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Theory of mind may have spontaneously emerged in large language models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.564499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.564499Z digest=sha256:df7fff80f3c1095a4f204faa8c699c8d74e0261c34448361346594d14a834868

Observation 10c98bb8-a181-4f77-b23d-126e62752d1b · outbound

This paper cites Breaking focus: Contextual distraction curse in large language models.arXiv preprint arXiv:2502.01609, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Breaking focus: Contextual distraction curse in large language models.arXiv preprint arXiv:2502.01609, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.658963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.658963Z digest=sha256:72bc8ed9affbf9df94b5860c924c32e7524b783a8b73f63b7f4ea7dff6c050e3

Observation 376a9404-d10c-4302-84ff-918afe3a131b · outbound

This paper cites FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.777761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.777761Z digest=sha256:89fe9830a4ab6cb07baa73460a184db3f3d3e6ecc46becaf26390373b2b1cf3b

Observation 834c773b-beb2-451b-93fe-ddf686541c21 · outbound

This paper cites Tomvalley: Evaluating the theory of mind reasoning of llms in realistic social context.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Tomvalley: Evaluating the theory of mind reasoning of llms in realistic social context

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.867408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.867408Z digest=sha256:8bc40ccbf07da49c7500931941a05fe4eb2f71a28664c60194aba2215aa895be

Observation 024dd819-cad6-4501-8a0a-ebaa415fc6e6 · outbound

This paper cites An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models An Open Review of OpenReview: A Critical Analysis of the Machine Learning Conference Review Process

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:18.940764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:18.940764Z digest=sha256:fff4a1e16c59568797d571e88e26bf515ceb919fec2b35d716983fdb7de9b29a

Observation 760e7e7f-9606-494e-b50d-4690a4b86426 · outbound

This paper cites The Open Review-Based (ORB) dataset: Towards Automatic Assessment of Scientific Papers and Experiment Proposals in High-Energy Physics.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models The Open Review-Based (ORB) dataset: Towards Automatic Assessment of Scientific Papers and Experiment Proposals in High-Energy Physics

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.218520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:45:19.052441Z digest=sha256:998d9c08fc9f0f3424323d5e2b8ab1dae3914bf6b5863e331fd48c280c8574ca

Observation 5f241ba2-4413-4f28-bf6e-d4d71f446b19 · outbound

This paper cites Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Superhuman ai for multiplayer poker.Science, 365(6456):885–890, 2019

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.122093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.122093Z digest=sha256:4afa2236e21b7e73081bcf8bf4888e3a73dcbc426ca5dafd48090d2507c60ec5

Observation f76f0b62-ae4c-48e7-a4df-c9c64a52d911 · outbound

This paper cites Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework.arXiv preprint arXiv:2502.13759, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework.arXiv preprint arXiv:2502.13759, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.216652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.216652Z digest=sha256:65e07e12322b0e4c530cff198f6f4f57efd6cc95edc25ae8c313503c1fc17dd0

Observation 49ecd989-94bc-4019-b6cf-4a2b8c9a7f91 · outbound

This paper cites Word form matters: Llms’ semantic reconstruction under typoglycemia, 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Word form matters: Llms’ semantic reconstruction under typoglycemia, 2025

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.323926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.323926Z digest=sha256:5690d29a1cef50e1915fd4200db6980b4d1f65d3b9f85786f9b7f5b6b89041df

Observation 3d23c6e0-dd3b-4975-9538-84b8fd3bba74 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.444402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.444402Z digest=sha256:5e5cefd303715ea9305dc97ba48a4affd83ba7be64923d2e83e949f3ff558794

Observation 6d0967d5-4ab6-42ab-9711-1d6e63576fe6 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models TrustLLM: Trustworthiness in Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.569117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.569117Z digest=sha256:8c07be7b3617393da04cdc35ade3b7eedb9a3029cce5c99807b6a16850a5cb40

Observation 0bf59a51-b52f-47b5-8970-cf99e84bf206 · outbound

This paper cites GPT-4o System Card.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models GPT-4o System Card

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.686427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.686427Z digest=sha256:d7f6f568461947f7a1294b49914721b73406a95d3f1a076934d8b7cd7fd9023d

Observation ab8e9951-1361-4cbe-a916-4c539efab799 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.https://openai.com/index/ gpt-4o-mini-advancing-\cost-efficient-intelligence/, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gpt-4o mini: Advancing cost-efficient intelligence.https://openai.com/index/ gpt-4o-mini-advancing-\cost-efficient-intelligence/, 2024

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.775775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.775775Z digest=sha256:7e49c632e041db8c0b1a76fdd416bfd027b3dff9993e65dcbe4689de48871c76

Observation 3fb4e4bc-8c02-4a0c-8ddf-2ea039558105 · outbound

This paper cites o3-mini.https://docsbot.ai/models/o3-mini, January 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models o3-mini.https://docsbot.ai/models/o3-mini, January 2025

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.864247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.864247Z digest=sha256:69c6f522a0edaa7481f4d79e61ca25df4b1e01efdc2dac3156d9ac7e2847f6a1

Observation 06751ed8-7c43-42b4-80ae-1ac0f08b4b78 · outbound

This paper cites OpenAI o1 System Card.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models OpenAI o1 System Card

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.032059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.032059Z digest=sha256:de8c765f6fc6701e2747e48dcc60de811567a323c7badca006fac9e64000c906

Observation 3ec7c936-bbd4-42c2-b9a7-5b26e2a44b73 · outbound

This paper cites Gemini 2.5 pro experimental.https://ai.google.dev/gemini-api/ docs/models, March 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Gemini 2.5 pro experimental.https://ai.google.dev/gemini-api/ docs/models, March 2025

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.213072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.213072Z digest=sha256:7a319deced9bc8c4baee496164cd215d4889b99309abc990de93eaacd17eae74

Observation caaf9613-11be-4b6b-a4c8-cd19e9351c51 · outbound

This paper cites Phi-4 Technical Report.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Phi-4 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.381718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.381718Z digest=sha256:de831c8ad7b12e6f72b540f698c95a2251dc6cb021cc6bd9fbc769131d837e43

Observation 212e4679-bfb0-4684-89a4-2ab9909d1e2c · outbound

This paper cites Llama 3.1-8b.https://huggingface.co/meta-llama/Llama-3.1-8B, 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llama 3.1-8b.https://huggingface.co/meta-llama/Llama-3.1-8B, 2024

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.477779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.477779Z digest=sha256:1f8ff0ebe90afc57a692781d2ac348c59f099bfb12db69bda85ae5573533f631

Observation 6c4a0d9d-9bee-408b-b4f2-5047eab5104b · outbound

This paper cites Llama 3.3-70b.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Llama 3.3-70b

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.590936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.590936Z digest=sha256:f70ae3d15ebc77f744502e282e9172153b81bae3976e9c0162124a543988b920

Observation 933701dc-3182-4c36-9915-900439b7e4c9 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.676823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.676823Z digest=sha256:60990660160d8ab726c426ca83cb5369267aef613b3e41818cdd7590af3fa6f4

Observation 66236904-c4d5-445f-8756-cb64ae00842c · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.773301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.773301Z digest=sha256:e88d529e9e54f9b893d6d9a796e1e8d050856a100e730c00a3d2dc9c407141d1

Observation ea100ba9-af50-4419-b410-edca7141a3cd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:20.998512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:20.998512Z digest=sha256:8cfa4ddf20169eb0a6d7a8bfb7f1903cdce896214301c6746ca7896a7b6adc1c

Observation 56e6ff9c-e1ef-4bbc-a08e-b14f0fb68dc8 · outbound

This paper cites an unresolved cited work.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:21.224254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:21.224254Z digest=sha256:7377bc573ea10e2788e8d04148b697a094cf4d10e899da911659c38767f423e3

Observation 3f2ed3cf-d633-4a12-b608-799c816a1980 · outbound

This paper cites 𝑢 is Criminal.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models 𝑢 is Criminal

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:21.292902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:21.292902Z digest=sha256:ea30bb47f03e686dee5ba4b5847ab08094bf662daa0c6cda2787ae79218fe0ca

Observation bbdad187-d187-4f70-a823-edddd76df300 · outbound

This paper cites Player𝑣 says Player𝑢 is the criminal.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Player𝑣 says Player𝑢 is the criminal

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:21.376172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:21.376172Z digest=sha256:5165ad138d01a1c6290e9423c1e1ad6837de404b978a0aa114a6f741c7a32379

Pith citing papers

Observation bfb07b62-2001-4fcf-bbaa-f2081bb94580 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.696600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:85ae3c2a4ae3ce8c2da5ecab03a3d07160e48291427d289913cacc3a2b15f6ee

Observation af7acf9e-6748-4e34-83d5-a80c1cd97551 · inbound

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs cites this paper.

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:46:29.084004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:37:45.062718Z digest=sha256:cab693795cb8b9f4fd8f7529e2b0f56b929b8ea7b10ca9f0dd83e10a9983c82d

Observation fde092a4-5309-4edc-9509-25b25cc7dd86 · inbound

OpenSkill: Open-World Self-Evolution for LLM Agents cites this paper.

OpenSkill: Open-World Self-Evolution for LLM Agents SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:56:59.840420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T00:47:38.201358Z digest=sha256:5e2b3b8d2ed2d0ddd3cd4b1b15bbb99e404ed55084f7573770f23b5447be505e

Observation 0785db62-f8a9-470a-88ea-c06e0644d9b4 · inbound

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents cites this paper.

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.390518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T07:04:51.970049Z digest=sha256:518a4bd74b906463f48a11e4fb681bd8fc2e3c4920f4ae1e11e1b57e7aac761c

Observation 650c9ba9-2d06-45ee-84df-fb206940328e · inbound

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models cites this paper.

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T10:39:44.819288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:42:26.713614Z digest=sha256:99bda5c5b49d3fd354335c5eb57871a5acbdd47ccfd2219cc3cc358495583a8d

Observation 24a806cd-70cc-405e-9802-a554c05f7c9b · inbound

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias cites this paper.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 218

Resolution
unresolved
no resolver link, observed 2026-07-14T02:33:34.084111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T02:33:34.084111Z digest=sha256:d180d0bd43a4ef1e05cd61791e0537237eba1d3982698d58ac2edf8307e652e4

Observation 7a631bd5-d10f-4067-a659-f09cd98f4aea · inbound

No One Wins in Nuclear War: A Social Simulation of Military Decision-making cites this paper.

No One Wins in Nuclear War: A Social Simulation of Military Decision-making SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:48.882711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:16:48.882711Z digest=sha256:70f124bdfed0ceb63d77d1c1c495b2fdd474cea3c3f15a4de4b3542aa2bce2ef